Voice is the hardest interface to automate well and the one with the least tolerance for error. Text gives a reader time to notice a poor answer and rephrase. A phone call gives them a pause, and a pause on a call reads as incompetence within about a second.
That does not make voice agents a bad idea. It makes the scoping decision considerably more consequential than it is for a chat interface.
Where they genuinely work
Structured, transactional calls with a narrow purpose: confirming an appointment, taking a meter reading, checking an order status, capturing a callback request outside working hours. These have a defined shape, a small vocabulary and a clear success condition, and handling them well is a real improvement on a queue or a voicemail nobody returns.
Out-of-hours coverage is the most defensible case. The comparison there is not against a person; it is against an unanswered phone, and almost anything beats that provided the escalation path is honest about what happens next.
Where they do not
Emotionally loaded calls, complaints, anything where the caller is already frustrated, and anything requiring judgement about an exception. Those calls need a person, and routing them to a system that cannot help produces a worse outcome than a longer hold time would have.
Regulated conversations belong in the same category. Where an answer could be construed as advice, the system should be gathering information and handing over rather than responding, and that boundary needs to be set explicitly before launch.
The technology is good enough for a narrow set of calls and nowhere near ready for the difficult ones. Most disappointment comes from deploying it across everything because it handled the easy calls well.
Sam Ortiz, Director of Engineering, Engineered With AI
Latency is the whole experience
A delay that would be unnoticed in chat is conspicuous on a call. Human conversation runs on very short gaps, and a system that pauses for a second before every reply feels wrong even when the answers are correct. This is why voice deployments frequently fail on infrastructure rather than on language.
Interruption handling matters just as much. People talk over each other constantly, and a system that cannot be interrupted, or that loses its place when it is, produces exactly the experience callers complain about.
Tell people what they are talking to
The agent should identify itself at the start. Callers work it out quickly regardless, and the ones who feel misled are considerably more annoyed than the ones who were told. It also sets expectations sensibly, so people ask simpler questions and reach for the human option sooner when they need it.
Design the handover before the conversation
When the agent hands over, the person receiving the call should get what has already been said rather than starting again. Making a caller repeat everything is the moment goodwill is lost, and it is a systems integration problem rather than a limitation of the voice technology.
What to measure
- Containment rate alongside whether the caller rang back within a day.
- How often callers ask for a person, and at what point in the call.
- Abandonment during the automated portion.
- Sampled recordings reviewed by a human weekly, particularly the short calls.
Short calls are where the failures hide. A thirty-second call that ended without resolution looks efficient in a dashboard and is usually somebody hanging up.
A sensible starting point
Out-of-hours callback capture, or appointment confirmation, with immediate escalation on anything unexpected. Prove it on that, listen to the recordings, and widen only where the evidence supports it. Our sales agent index covers the inbound enquiry side by sector.
Accents, noise and the practical failure modes
Speech recognition performs unevenly across accents, background noise and poor line quality, and those are ordinary conditions rather than edge cases. A system tested by its builders in a quiet office will behave differently on a call from a van on a motorway, which is precisely where a trade customer is calling from.
Test with real recordings from your actual call mix before committing. Where recognition is unreliable for a meaningful share of callers, the honest answer is a narrower deployment or none, because a system that mishears a name or an address produces work rather than saving it.
Considering a voice agent?
We will tell you honestly which of your call types suit it and which do not, before anything is built.





