Voice AI · · 6 min read
What a voice AI agent actually has to get right
Latency, turn-taking, phone-line audio, handoff and the metrics that matter — the engineering realities behind a voice agent people don’t hang up on.
By XYNEXIA
Most voice AI demos are recorded in a quiet room, over a good microphone, by someone who knows what the agent expects. Real calls are none of those things. The gap between the two is where voice projects succeed or fail — and very little of it is about which language model you pick.
Latency is the product
In conversation, people expect a reply within a fraction of a second of finishing a sentence. Much longer and the caller starts repeating themselves, talking over the agent, or assuming the line has dropped.
A voice agent is a pipeline — speech recognition, reasoning, speech synthesis, plus telephony on both ends — and every stage adds delay. The practical consequences:
- Stream everything. Transcribe while the caller is still talking, start generating as soon as intent is clear, and start speaking before the full reply is written.
- Measure latency per turn, at the 90th and 95th percentile — not the average. Callers remember the slow turns.
- Keep tool calls off the critical path where possible, or cover them naturally (“Let me check that for you”).
Turn-taking is harder than understanding
Knowing when a person has finished speaking is a genuinely hard problem. Pause too long before replying and the agent feels slow; reply too quickly and it interrupts people who were only thinking.
Callers also interrupt — to correct a detail, to say “no, not that one”, or simply because they already know the answer. An agent that can’t be interrupted, or that loses track of what it had already said when it is, feels robotic immediately. Handling barge-in well is one of the clearest signals of quality.
The phone line is a hostile environment
Traditional phone audio is narrowband and compressed. Add traffic, speakerphones, background conversations, accents and code-switching between languages, and recognition accuracy drops sharply compared with the demo.
Test with real recordings from your own callers, and design the conversation to be robust: confirm critical details (names, numbers, dates, amounts) by reading them back, and prefer questions with a small set of likely answers when accuracy matters.
Handoff is a feature, not a failure
Every voice agent will meet calls it shouldn’t handle: complaints, edge cases, emotional situations, requests outside its scope. The goal isn’t to keep those callers talking to AI. It’s to recognise them early and transfer to a person with the context already captured, so the caller never has to repeat themselves.
Measure outcomes, not containment
“Containment rate” — the share of calls that never reach a human — is easy to game. An agent that frustrates callers into hanging up has excellent containment. Better measures:
- Task completion: did the booking, update or answer actually happen?
- Escalation quality: when it handed off, was it the right call, and was context passed on?
- Turn latency at p90/p95, and interruption handling.
- Listening. A weekly review of real calls finds problems no dashboard shows.
None of this is glamorous. But it’s the difference between a voice agent people tolerate and one they’re happy to talk to.