The Awkward Pause Was Your Product Problem. It Now Costs $3 an Hour.
Voice agents felt fake because of turn-taking, not intelligence. Now that full-duplex speech is a rented API call, the hard part of your voice product moved somewhere less fun
I've sat in two different companies where a team spent months on the pause.
Not on what the agent said. On the half-second of dead air after the caller stopped talking, before the bot started. That gap is what made everything sound like an answering machine from 2011, and every founder who heard the demo said some version of "it feels off" without being able to name why. So the engineers went after it. Buffer tuning, voice activity detection, endpointing thresholds, barge-in handling so the caller could interrupt without the whole thing falling over. Real work. Genuinely hard. And roughly nine months of combined engineering time that a vendor then shipped for free in a point release.
That's the pattern I want you to notice, because it just happened again at a scale that should change your budget.
OpenAI launched GPT-Live-1 this week at $0.05 per minute. Full duplex, meaning the model listens and speaks at the same time instead of taking turns. It handles interruptions. It makes acknowledgement noises while it's thinking. Early tests report 80% fewer awkward interruptions than turn-based systems. Twelve voices, tone and pacing steerable from the system prompt.
Three dollars an hour for the thing that made your product feel fake.
The intelligence was never the problem
Here's the part founders got wrong for two years. Everyone assumed voice agents sounded robotic because the model wasn't smart enough. It wasn't that. The text coming out of a 2024 model was already better than half the offshore call centers it was competing with. The problem was mechanical: humans overlap when they speak. We say "mhm" in the middle of your sentence. We cut you off when we already know the answer. We start talking before we've finished thinking.
A turn-based system can't do any of that, and your brain clocks it in under two seconds. Not as "this AI is dumb." As "this is not a person and I don't want to be on this call."
So the entire perceived quality of your voice product was gated on a plumbing problem, and the plumbing problem is now rented. That's the good news. The bad news is what it does to your plan.
Stop budgeting for a voice layer
I'll say this plainly because I think a lot of money is about to go into the wrong line item. Do not hire a speech or audio ML person for this. Not this quarter, not next.
Whatever you build in-house to manage audio streams, latency, and interruption will be worse than the vendor's version within six months and irrelevant within twelve. This layer is going to get commoditized again by Q2, the same way transcription did, the same way text-to-speech did. Google, Anthropic, and four well-funded startups will all ship full duplex, and the price will fall. Building there is building on a floor that keeps dropping.
If you've already got a quote for a voice infrastructure build, that quote is for a thing you will rent for pennies before you finish paying for it.
The counter-argument I take seriously: some businesses have genuine audio constraints, like medical dictation in noisy rooms or telephony over bad PSTN lines where vendor models degrade. If that's you, you'll know, because you'll have real recordings that fail and you can measure it. Everyone else is imagining a constraint to justify a hire.
The part nobody demos
Every voice demo ends when the call ends. That's the tell. The demo is engineered to stop at the exact moment the actual product would start.
Think about what your business needs from a phone call. A patient calls to book an appointment. Great conversation, agent sounds human, very impressive. Now: did the appointment get into the scheduling system? Under the right patient record or a duplicate? What happens when the requested slot got taken forty seconds ago by the front desk? Who gets notified when the agent couldn't find a match? What does the human see tomorrow morning when the patient shows up and the slot doesn't exist?
None of that is voice. All of it is the product.
I build MVPs for founders now, and this is the layer I'd put an engineer on and refuse to rent. Three things specifically.
The workflow. What the agent is allowed to decide on its own, what it must escalate, and where the handoff to a human actually lands. This is a business rules problem wearing a technical costume, and you are better qualified to define it than any engineer you hire.
The integrations. Your CRM, your calendar, your billing, your case management system. This is unglamorous, slow, full of bad APIs and undocumented edge cases, and it is where the value lives. It's also the thing a generic voice vendor will never do for your specific stack.
The record. What the agent did, why, what it heard, what it wrote, and what it got wrong. Not call recordings. A structured, queryable trail of actions taken. You need this for disputes, for compliance, for firing the agent from tasks it keeps botching, and for the conversation with your first enterprise customer who asks what happens when it's wrong. I've watched teams treat this as a phase two, and phase two arrives in the form of an angry customer and nobody able to reconstruct what happened.
What I'd actually do with the money
If you run a phone-heavy business (intake, scheduling, qualification, research calls) you can ship a credible voice MVP this quarter without a single speech engineer. Rent the voice. Buy the workflow.
The honest scoping exercise takes an afternoon: list the twenty most common calls you get, mark which ones end in a system write, and count how many distinct systems are involved. That number, not the audio quality, tells you what your build actually costs. It's the same exercise I run at the start of a fixed-scope MVP at $14,500, and it's usually the first time a founder sees that their voice product is 15% voice.
The pause is solved. The part where the call turns into a booked appointment, a qualified lead, or a closed ticket is still yours to build, and it always was.
If you're staring at a voice roadmap and can't tell which half to pay for, grab 30 minutes and I'll tell you which half I'd cut.