Back to all posts
Technical StrategyInterim CTO

24% Alone. 82% With a Human. The Gap Wasn't Coding.

A new benchmark shows the same frontier model triples its success rate when paired with someone who knows the business, and what that means for your next hire

I've been dropped into roughly ten companies as a fractional or interim CTO. Different industries, different stages, different reasons for calling me. The first task has been identical every single time.

Nobody wrote down what the thing was supposed to do.

Not in a document, not in a ticket, not in anyone's head in one piece. It exists in fragments. The founder knows what they promised the first three customers. Ops knows which of those promises is quietly being fulfilled by a person copying rows into a spreadsheet. Support knows the four edge cases that generate most of the angry emails. Nobody has ever assembled these into one picture, because assembling it was never anyone's job.

So I spend my first two weeks doing archaeology. Reading Slack backwards. Sitting in on support calls. Asking the same question to five people and mapping where the answers diverge. That output, the reconstructed spec, is worth more than anything I write afterward.

I've always struggled to explain to founders why that phase is the expensive part. This week someone put a number on it.

The benchmark

Sierra published Hyper-๐œ-bench, which is a fairly clever setup. They drop a developer agent into a sandboxed workspace containing the records of a simulated business. It has a simulated client it can message at any time. Its job is to recover what the business actually needs, design the thing, and build a working customer-service agent that has to run inside a cost budget per conversation.

In other words, the benchmark is testing the whole job, not the coding part of the job.

Claude Opus 5 with max reasoning, running in Claude Code, passed 23.9% of the held-out evaluation tasks working alone. The same class of model paired with an engineer who had deep context reached 82.2% on the same tasks.

Same model. Same tools. Same problem. 3.4x difference.

I want to be careful here. This is one benchmark from a company that sells agent infrastructure, and I haven't run it myself. But the shape of the result matches what I've watched happen in real companies for years, which is why I trust it more than I'd trust a number that surprised me.

What the model was actually missing

Read the setup again. The agent had access to the business records. It could message the client whenever it wanted. It had every technical capability required.

It still failed three quarters of the time.

It didn't fail because it couldn't write code. Writing the code was the part it was best at. It failed because it didn't know which of the twelve things in the records mattered, which client answer was the real requirement versus a passing preference, and which unstated rule would blow up in production. It couldn't tell the difference between what the business said and what the business meant.

That gap is exactly what a person with context fills. Not by typing faster. By knowing which question to ask on day one instead of day nine, and by recognizing a wrong answer before it becomes 4,000 lines of confidently wrong architecture.

I have a specific memory of this. At Lucky Day we were scaling toward 10M monthly active users and someone proposed a change to how we handled a particular class of user session. Technically sound. Reviewed cleanly. It would also have quietly broken a revenue path that existed for historical reasons nobody working on that team had lived through. I knew because I'd been there when we built it. That knowledge wasn't in the code, wasn't in the docs, and would not have been recoverable from the repository by any reader, human or otherwise.

Every company has dozens of those. They're the difference between 24% and 82%.

What this changes for your budget

Most non-technical founders I talk to are running one of two mental models right now.

The first: AI can just build it, so I need a cheap builder or maybe no builder at all. The second: I need to hire a CTO, and I can't afford one, so I'm stuck.

Both are budgeting for the wrong line item. They're both asking who writes the code. The benchmark says that isn't the variable. The variable is who holds the context and transfers it.

So the question you should actually be pricing is: who in my company can sit down and produce a complete, honest, contradiction-free description of what we're building and why, including the parts that are ugly?

If the answer is you, and you can do it in writing, then hiring a strong builder or pointing good tooling at it may genuinely work. You are the context holder. That's a real role and it's the scarce one.

If the answer is nobody, no amount of model capability rescues you. You will get an impressive artifact that does not survive contact with your actual customers, and you will find out four months in. I've been called to clean up exactly that, more than once. The rebuild is not the expensive part. The four months are.

The uncomfortable version

Here's the part founders don't love hearing. The reason nobody has written the spec is usually that writing it would force a decision the company has been avoiding.

Which customer segment are we actually serving. What happens when those two rules conflict. Are we willing to say no to the biggest logo's weird requirement. The spec doesn't exist because it's a document full of decisions, and decisions are harder than tickets.

So when I do that archaeology in the first two weeks, half the value isn't discovery. It's forcing the conversation that produces the answer. Once that exists, the building genuinely is faster and cheaper than it has ever been. That part of the hype is real.

If you want someone to do that reconstruction with you, that's most of what a two-week fractional CTO sprint is: $4,500 to walk out with the document your team has been working around for a year. If the spec already exists and you just need it built, a fixed-scope MVP is $14,500, and it's cheap precisely because the hard input is already there.

Stop budgeting for who writes the code. The model writes the code. Budget for the person who knows what it's supposed to do, and notice how few of them you have.

Written by Hootan Nikbakht

Building something?

Get a second opinionbefore it costs you a quarter

I build MVPs for non-technical founders and step in as interim CTO when there's a leadership gap. If this one hit close to home, bring me the repo.

Or email hootan@nikfam.ai.