Why AI voice agents fail in production is not a question demos can answer, because demos never run long enough to fail. We have run AI voice agents on live medical phone lines for a year: real patients, real diaries, real money, across a network of 40 doctors. This is the list of what went wrong, what was actually causing it, and what fixed it. Almost none of it is in any vendor's documentation. A few items contradict the documentation. One of them caused a production outage that we caused ourselves.
It is written for the engineer or technical owner who is about to ship a voice agent and does not yet know what they are walking into. Practices are described by specialty only. No patient data, phone number, address or identifier appears anywhere in it.
| Metric | What a year in production showed us |
|---|---|
| Real side effects from one test session | ~277 bookings, ~300 emails, ~270 payment links |
| Tests giving the same verdict across 3 runs | 54 of 72 |
| Message-taking calls with a false "I've logged that" | About 1 in 21 |
| Long pauses caused by silent webhook calls | 39% of 3 to 4 second pauses on one line |
| Recurring defects that were model-level | None. All prompt-level or config-level |
The One Principle Underneath Most of This
A successful write is not correct behaviour. An API returning 200. A config that reads back exactly as you set it. A test suite flag saying passed. An agent saying "I've logged that for you." None of these prove the thing works.
Every serious incident we have had came with a clean-looking confirmation sitting in front of it. Behaviour is only proven by the transcript, the tool-call record, the delivered email, the carrier call record, or a real phone call. If you take one thing from this post, take that. Most of the 36 lessons below are a specific instance of it.
The agent says: I've logged that for you.
The platform marks the call successful.
The API returns 200 and the config reads back as set.
The suite says passed.
The agent works in the dashboard widget.
A 200 response, a config that reads back correctly, a passing test flag or the agent saying it logged something are all signals, not proof. Behaviour is proven only by the transcript, the tool-call record, the delivered email, the carrier call record or a real phone call.
What We Run, and Why We Built It on ElevenLabs
The agent is the visible part. Around it we built a whole ecosystem, and most of the lessons in this post come from the seams between the agent and that ecosystem rather than from the agent itself.
Veda, the voice platform behind our 40-doctor medical voice assistant case study, is our own engineering wrapped around a managed conversation layer:
- A practice management system (PMS) that holds doctors, patients, calendars and bookings in one governed database, with each doctor's records isolated from every other doctor's.
- Custom doctor schedule overrides, because a real clinic diary is never just "opening hours": a doctor adds a clinic, drops a session, moves a list to another site.
- A custom doctor forms system, so the information a doctor needs before a consultation is collected by the practice's own forms rather than improvised in conversation.
- A payment system with two routes: a secure pay-by-link sent by SMS, and in-call card entry with DTMF masking, where keypad tones are replaced with a flat tone before they reach the AI, the transcript or the recording.
- The conversation layer on ElevenLabs Agents (formerly ElevenLabs Conversational AI), connected to all of the above through tools.
A patient call reaches the conversation layer on ElevenLabs Agents. Through tools, the agent reads and writes the practice management system, which holds doctors, patients, calendars and bookings, applies custom doctor schedule overrides, uses the custom doctor forms system, and hands payment to two routes, a secure pay-by-link by SMS or in-call card entry with DTMF masking.
Why the platform choice mattered more than the model
The underlying voice platform is one of the most consequential decisions in a voice AI project, and it matters more than which language model you pick. That sounds like it contradicts Lesson 33 below, where model choice made almost no difference. It does not. The model is one swappable component. The platform is everything you build on for the life of the project: the telephony bindings, the tool system, the turn-taking controls, the testing framework, the API you automate against.
We chose ElevenLabs for reasons that held up over the year:
- Documented APIs for everything. Agents, tools, tests and phone numbers can all be created and changed through the API, which is what makes the fleet audits in Lesson 35 possible. You cannot assert what you cannot read.
- Leading speech models. Voice quality and speech recognition are the parts callers judge first.
- A real testing suite. Agent Testing supports full simulated conversations, next-reply tests and tool-call tests, with tool mocking, a CLI, and a
repeat_countoption that runs a test several times and reports a pass rate. Several lessons below are about how to use that suite well. None of them would exist if the suite did not exist. - The tooling around the agent: system tools such as skip turn and end call, call transfer, and a steady pace of development.
Two things weighed against the alternatives. Some newer platforms lead with benchmark scores that cannot be independently verified. A latency or accuracy figure measured by the vendor, on the vendor's test set, tells you very little about your callers on your phone lines. And open-source stacks vary enormously depending on where and how you host them. Latency, speech recognition and language recognition all move with the hardware, the region and the serving setup, and you then own the whole ecosystem around the agent as well: telephony, testing, monitoring, transfers. That is the right call for some deployments, which our voice AI development guide covers, but it is a bigger build than it looks.
None of the platform traps in Lesson 32 changed that decision. Every platform has edges. The ones worth building on are the ones where you can find the edge through the API and engineer around it.
Part 1: Testing a Voice Agent Without Hurting Anyone
Lesson 1: Our tests took patient message delivery down for a day
This is our worst incident, and the most useful thing we can tell anyone.
A model-benchmarking session ran roughly 37 simulation test suites with tool mocking turned off. Those suites booked about 277 real calendar appointments, sent about 300 real emails and created about 270 real payment links. The sends exhausted the email provider's daily per-account quota. That quota is shared with live patient traffic, so real patient message delivery stopped for every practice on the platform until the 24-hour window rolled over.
The test data made it worse. It used realistic names, so once it landed in inboxes and diaries it was indistinguishable from genuine patient bookings and could not be safely cleaned up.
Four rules came out of it, and we have not broken them since:
- Mock anything with an outward-facing side effect: every send, booking, payment link and SMS. The real send is the deliberate exception, not the default.
- If test data can reach a human, the word TEST must be in it. In the patient name, the email subject, the calendar title, the SMS body and the payment description. Never a plausible-looking name. A realistic name in a diary is unrecoverable.
- Send the fewest real messages that prove the thing. Once per behaviour, not once per assertion. Never loop a send, never send per parametrised case. Ten template variations are ten offline renders and at most one real send.
- Run only the tests for the issue in front of you. A broad suite after a narrow change produces noise that buries the signal you wanted.
A healthy shape is roughly 200 offline tests and one deliberate end-to-end send. If a change needs dozens of real messages before you believe it, the test design is wrong, not the quota. This is the sharp end of a rule we give every client, voice or not: sandbox every system with real-world consequences before an agent touches it.
Lesson 2: The dangerous mocking setting is the one that looks careful
ElevenLabs offers three mocking modes for simulation tests: mock none, mock all, and mock selected tools. Everyone reaches for selected, because it looks deliberate.
Auditing 85 stored tests across two agents, 40 would have fired a write tool for real. Eleven were set to none. The other 29 were set to selected, with the practice's own message-sending tool simply missing from the list, including most of a suite written three weeks earlier by someone who believed they had mocked it.
- Safe: write tools mocked53%45
- Set to 'selected', send tool missing from the list34%29
- Set to 'none'13%11
A write tool that is absent from a selected list runs for real, and nothing in the test name or the run output tells you. Mock all, plus explicit overrides where you genuinely need a real call, cannot silently miss one. Delete "selected" from your team's vocabulary.
Lesson 3: Three quieter mocking traps
- A per-call mocking override was silently ignored on our setup. The stored test config always won. You have to fix the stored test, not the run.
- A mocked write tool with no stubbed success response errors, and the agent's retry behaviour then produces failures that look exactly like behaviour bugs.
- A test inbox is not an exemption. The send still draws down the shared account quota that live patients depend on.
Lesson 4: Audit the whole suite, not the test you just wrote
New tests being correct tells you nothing about the ninety already stored. Walk every test and assert its mocking mode whenever you add an agent, add a tool, or inherit a suite from someone else.
Lesson 5: Your test suite is probably lying to you about effect sizes
We ran the same 72-test suite three times against three configurations at temperature 0.2.
| Comparison | Tests that changed verdict |
|---|---|
| Run A vs Run B | 15 |
| Run A vs Run C | 9 |
| Run B vs Run C | 12 |
Only 54 of 72 tests gave the same verdict in all three runs. Only two failed in all three, and those two were the only real defects in the list. The configurations scored 63, 62 and 62. We nearly reported a difference. There was none: the noise floor of the suite was larger than the effect we were trying to measure.
This is not peculiar to voice. Thinking Machines Lab showed in September 2025 that even at temperature 0, 1,000 runs of the same prompt on an open model produced 80 different completions, because server load changes batch size and most inference kernels are not batch-invariant. A voice test adds a simulated caller, who is also a model, on top of that. Run 3 to 5 rounds per arm before claiming anything. One round measures luck. ElevenLabs' repeat_count exists for exactly this, and we would treat any single-run "regression" you have ever reported as suspect. We had to go back and re-read ours.
Lesson 6: Beware a statistic that splits on something other than the variable
We found a 92% versus 79% pass-rate split and briefly believed it proved a feature worked. It did not. The two groups were different tests. Controlling for that, the difference vanished.
Lesson 7: Classify failures by kind before touching the prompt
Quota errors, infrastructure timeouts, a scenario that never triggered, and a stale test all look like behaviour failures in a results table. In one of our runs, 29 "regressions" turned out to be an exhausted credit balance. Sort the failures into infrastructure, test-design and behaviour buckets first. Only the last bucket is a prompt problem.
Lesson 8: A judge-criteria field that silently ignored our edits
Stored tests on our platform carried both a singular success_condition string and a plural success_conditions list. The grader read the plural one. Writing only the singular returned 200, read back looking perfect, and changed nothing.
We "tightened" four tests, watched them keep failing with rationales quoting the old wording, and spent two days blaming the grader for being over-literal. The edits had never landed. The principle at the top of this post, in one bug.
Lesson 9: Simulations cannot test anything with audio timing in it
Text-turn simulations cannot test latency, interruptions, or a caller reading a phone number in chunks: both sides simply wait and the run times out. Turn-taking and voice settings scored identically across a whole suite whether we set the agent to patient or eager. If a defect lives in timing, no amount of simulation will find it. You need a real call, which is why the supervised batch of real calls in our AI voice agents guide is not optional.
Part 2: Silence, Noise and Turn-Taking
Lesson 10: A punctuation-only turn is noise, and the right answer is silence
On one line, the most common "turn" in the transcripts was a turn containing only punctuation: ..., .., ?, a dash. That is not the caller speaking. It is the transcriber picking up a knock, a car, a television or an over-sensitive microphone. The caller has said nothing.
Our own agent template told the agent to answer it, with a line along the lines of "if the input is pure noise, say you didn't catch that and ask how you can help". That one line, inherited by every new agent we built, produced this on test: five noise artifacts in a row, five replies, each escalating ("take your time", "are you still there?", "just checking in"), and then the agent closed the call on a caller who was sitting right there.
The fix: a punctuation-only turn means say nothing, using the platform's stay-silent action (on ElevenLabs, the skip turn system tool). If no such action is available to the agent, no prompt rule can fix this. You are telling the model not to speak without giving it anything else to do. And never let the agent emit ... or a fragment itself.
A punctuation-only turn is noise, and a filler alone such as mm or uh is not a turn either. Both mean say nothing and leave the silence counter where it is. Words like yeah, hello, right or sorry mean the caller is there and the agent should respond to them.
Lesson 11: Noise must not advance your silence counter
This is the rule that breaks most easily. Five artifacts in a row are zero turns from the caller, not five pauses. Do not answer the first, and do not "check in" on the third.
- A filler alone ("mm", "uh", "oh") is not a turn either. The caller is still thinking.
- But "Yeah", "Hello", "Right" and "Sorry?" mean the caller IS there. Answering those with a silence prompt is its own failure. We had to enumerate both lists explicitly in the prompt.
Lesson 12: Escalate, never repeat, and a few seconds of quiet is not your cue
A caller struggled to give their mobile number, and the agent asked "could you give me the number once more" with identical wording each time. Nothing reads more robotic, and it visibly frustrated them.
What works is a four-step ladder, plus a hard rule that no prompt sentence is ever used twice in one call:
- 01Say nothing
The caller is thinking, finding a letter or putting on glasses. Stay silent and keep waiting.
- 02A specific nudge
Tie it to whatever you are waiting for, such as the rest of the mobile number.
- 03Different wording
Use a sentence the agent has not used anywhere else in the call.
- 04Check the line once
Ask once whether the caller is still there. No prompt sentence is ever used twice in one call.
Resist nudging after three or four seconds. People pause to find a letter, put on their glasses, or think of a date. Nudging rushes them and makes them lose their place. It is the same lesson we learned earlier about eagerness, from a different direction: an agent tuned to respond fast feels broken, not fast.
Lesson 13: Filler-suppression settings only suppress barge-in
The platform has an "ignore these terms" list for interruptions. We added 24 filler terms and expected the agent to stop responding to "um". It kept responding.
The setting stops a filler from interrupting the agent mid-speech. It does nothing about the agent replying to a filler when it is the agent's turn. Those are two different problems, and only the second one is a prompt problem.
Worse, the list had a merge with platform defaults flag. One agent had the list empty and the merge flag off, which disabled the platform's own defaults too, so it barged in on every "um". An empty list is not neutral.
Part 3: Numbers, Names and Email Addresses
Lesson 14: Language models regenerate numbers, they do not copy them
Asked to repeat a number, one agent said it, then said it again in a different grouping, then again, and the loop ended a live call. Models do not copy a digit string. They regenerate it, and each attempt is a fresh chance to drift.
Two fixes, used together. Say numbers as hyphen-joined words in fixed groups ("zero-seven-eight-three-three, four-five-eight"), treating each group as text being copied rather than a number being recalculated. And say any given number at most once per call, without offering to repeat it.
Lesson 15: A correction replaces what came before, and a pause is not the end
A caller corrected themselves mid-number, and the agent stitched the abandoned fragment onto the correction, producing a 15-digit read-back. "Sorry", "no", "I mean", or a repeated opening group all mean: discard the fragment.
The opposite failure is cutting in. People read numbers in chunks. If what you have is shorter than a full number, the right move is to stay silent and wait.
Lesson 16: Count the digits and check the prefix before you read anything back
A too-short number was accepted and read back confidently, and the caller said "yes". A caller will happily confirm a number they did not hear properly, and their confirmation makes a wrong number look verified. Count the digits and check the prefix before accepting. Never pad with an invented leading zero to fix the count, and never trim. Ask for the missing part.
When the caller corrects themselves, the earlier fragment is discarded. If the digits so far are shorter than a full number, the agent stays silent and waits. Once the digit count and prefix check out, the agent reads the number back once as grouped words. If they do not, it asks for the missing part and never pads or trims.
Lesson 17: When you add a rule, delete the one it contradicts
One of our agents carried "after giving a number, offer to repeat it" immediately alongside the say-it-once rule that existed because a repeat loop had ended a live call.
Adding the right rule is half the job. Search for and delete what it contradicts. We hit this repeatedly: a "call the tool exactly once" rule suppressed the first call because the model believed it had already logged; old sections left behind after a redesign created a competing path and intermittent stray tool calls.
Lesson 18: A spoken read-back cannot catch a spelling mistake
"Alastair" and "Alistair" sound identical. A caller will confirm a read-back of a name you are about to record wrongly, and that is exactly how a team ends up unable to find callers in its own system.
You need both: the spoken read-back confirms you heard the right person, and only the spelling confirms you will record the right name. And once a caller has spelled it, the spelling is the truth. Hear "Hallworth", get spelled "H-O-L-W-O-R-T-H", and you record Holworth. Recording the first hearing anyway defeats the entire point of asking.
Lesson 19: The etiquette of spelling, each rule learned by breaking it
- Never ask for a spelling the caller already gave. 26 occurrences in 30 days of calls on one line. It is the single most irritating thing in our transcripts.
- A full name is a first name and a surname. One word is not a full name.
- Accept a refusal to spell immediately. Never press twice.
- Reassemble a spelled name into a word. Say "Smith", never "S-M-I-T-H".
- Letters, not the NATO alphabet. Our clients consistently preferred it.
- Never silently auto-correct a name to something more familiar.
Lesson 20: Speech-recognition keyword boosting is the cheapest fix you are not using
A real transcript contains a consultant's name mangled into "Dr Rada. Mommy". Feeding the recogniser your actual consultant surnames, hospital names and local place names is the cheapest accuracy improvement available, and we left the keyword list empty on several agents for months.
Lesson 21: What the agent speaks and what it stores are different things
We told an agent to say "at" and "dot" aloud when reading an email address back. It then stored those words. A payment tool received an address like name99atyahoo.co.uk, which is dead, so the receipt went nowhere.
Any read-back rule must state explicitly that the spoken form is speech only and that the stored value keeps the real @ and .. It is the same root cause as number drift: the model regenerates rather than copies.
Part 4: Never Let the Model Decide What Time It Is
Lesson 22: A timezone field does not make an agent time-aware
At 08:53 one agent said "there's no one in the office" and then, in the same call, "our office is open now". It had the correct time. It got the comparison wrong seven minutes before nine, and a hard-coded script in the prompt disagreed with its own arithmetic.
- Setting a timezone on an agent only controls how an injected time is formatted. Writing "you know the current time" into a prompt does nothing. You must reference the platform's current-time variable in the prompt text itself. Any prompt that discusses "the current time" without containing that variable is not time-aware. We found four agents in that state.
- If behaviour depends on time, compute open or closed on your server and pass a boolean. Do not ask a language model to do time arithmetic.
- Search your prompts for hard-coded office-state claims. "No one is in the office at the moment" sitting inside a script ignores the clock entirely.
Lesson 23: Use wording that is true at every hour
State the office hours, plus "a member of the team will call you back during office hours". Avoid "shortly", "soon", "today", and "when the office is open", which callers hear as "it's closed now". Never let the agent name a specific callback day, and keep the bank-holiday list current.
Lesson 24: Check the timezone of every log you reason over
Our carrier's call-detail records store timestamps in UTC, not local time. Misreading them led us to "discover" a daylight-saving bug that did not exist, and to make a production routing change on the strength of it. The dashboard is a second trap: it renders timestamps in your browser's timezone while routing and prompts run in the service's timezone.
Part 5: Bookings, and the Agent That Says It Did Something
Lesson 25: Trust your own availability endpoint
Our scheduling endpoint returns bookable slots that already exclude existing appointments. Someone added a prompt rule telling the agent to double-check each offered slot against the raw calendar before confirming.
Some raw calendar entries came back with null start and end times, so they were not reliable for that comparison. The agent offered a slot, the patient accepted, the agent "checked", wrongly concluded it was taken, and said "that slot has just gone", then offered a different time from the same list it had just decided not to trust. On nearly every call.
The plausible-sounding reasoning ("availability only knows opening hours, so check the calendar first") was wrong, and writing it into the prompt mandated the bug. Identify your authoritative source, and do not let the agent second-guess it. Only a genuine conflict at booking time should override it. Our AI scheduling agents guide covers the constraint model behind this.
Lesson 26: "I'll log that now" is speech, not a tool call
About 1 in 21 message-taking calls confirmed logging to the caller and never called the tool. The platform's call-success flag marked those calls successful, and the auto-generated summary repeated the agent's false claim back to us.
- Never monitor with the platform's success flag or its generated summary. They restate the agent's own claims. Read the tool-call records. Our guide to AI monitoring in production covers the wider version of this.
- Search every promise in your prompt ("I'll pass this on", "I'll transfer you", "someone will call you") against the tools actually attached. A promise with no tool behind it is a bug, and from the caller's side it fails silently.
Lesson 27: Prose does not stop fabrication. Structure does
An invented date of birth survived three escalating warnings in the prompt. Use structural gates instead: a required tool parameter, a state flag that must be set, backend validation that rejects the write.
- An "absolute rule" block works until you have six of them. When everything is absolute, nothing is.
- Examples get taken literally. A "log it early" rule illustrated with a payment example made the agent start asking every payer for their name before offering payment options. When you add a rule, check every section it could touch.
Part 6: Telephony and Payments, Where the Silent Failures Live
Lesson 28: A call that never connects leaves nothing to debug
- An unregulated phone number drops calls instantly and leaves no call record at all. There is nothing to debug. Check the number's regulatory status first.
- "Live" means a phone number is actually bound to the agent. An agent that works perfectly in the dashboard test widget is not live.
- Respect provider rate limits. One request per collection. Never loop-probe endpoints or fan out a request per object.
Lesson 29: Transfers and routing rules fail where nobody is watching
Warm transfer is often carrier-specific. On ElevenLabs, the transfer documentation is clear that the warm message to the human operator works only through the native Twilio integration, and SIP REFER transfers need a trunk that allows them. On our carrier, the standards-based method failed outright with a carrier rejection code. We found that out after promising the feature. Verify transfers on your exact carrier before you sell them.
Date-bounding a time-of-day routing rule can silently kill the time window. We added a date range to an out-of-hours interval to cover bank holidays. It disabled the time-of-day behaviour entirely and took the agent off three phone lines for out-of-hours calls for a week before anyone noticed. Out-of-hours failures are invisible, because nobody is watching at those hours.
Lesson 30: Payment status codes will mislead you
- Key off the result field, never the HTTP status. On our payment provider's API,
201covers both authorized and refused, and202covers both sent for settlement and sent for cancellation. A naive "2xx means success" check books a refused payment as paid. - The auto-capture field differed between the provider's own API products. We used the field from the wrong product's documentation and got
400until we worked out which API our account was actually on.
Lesson 31: Payment invariants belong in the database, not in convention
- A new payment identifier on an already-paid session means you are charging twice. Make that an explicit invariant with a database-level guard.
- Do not expire payment records with a generic sweep job. Ours deleted the associated bookings.
- Receipts must send exactly once, on the transition into paid. What prevents a double send is claiming that transition inside the transaction, not a replay guard bolted on afterwards.
- A pay-by-phone transfer tool can block your entire test suite. Ours declares a parameter that only exists on real inbound calls, so while it was attached, every simulation on that agent failed before it started. We keep a sandbox agent with a parameter-free copy of the tool.
The payment design itself, both routes and why DTMF masking keeps the AI outside PCI DSS scope, is written up in our AI receptionist guide.
A compliance note: keep regulated data out of the model's reach
Every lesson in this part is easier when the most sensitive data never touches the AI in the first place. On these UK medical lines, patient data stays in the UK, which is how the platform meets UK GDPR by architecture rather than by policy; a data handling layer keeps patient identifiers out of AI processing; and DTMF masking means card numbers never reach the model, the transcript or the recording. Design that boundary before you tune a single prompt, because no amount of prompt work can make an agent safe to hold data it should never have received. Our guide to AI automation under UK GDPR and HIPAA covers what each regime allows an agent to do.
Part 7: Platform API Traps
Lesson 32: Changing one thing is rarely changing one thing
- Writing the system-tools array silently cleared the webhook-tools array, and the platform then deleted the orphaned webhook tool. This destroyed a live production tool twice before we understood it. Create the replacement first.
- A partial config write can reset sibling fields. Always re-read the agent afterwards and print the fields you did not intend to touch.
- A dashboard model change silently reset a per-tool speech setting, which caused a customer ticket.
- Documented enum values your account may not accept. One pacing setting is documented with three levels; our account accepted two. A reasoning-effort parameter is accepted by one model family and returns
400on another, so switching models means clearing it. - Documentation can be wrong. Twice we hit a mismatch: a recommended voice model rejected for English agents, and a documented default that did not match reality. Our source-of-truth order is now the live API and its machine-readable spec first, live vendor docs second, dated blog posts last.
- Internal phone numbers must never enter a prompt. The model will eventually read one out.
- Transcripts use curly apostrophes. A search for
I'll logwith a straight quote found nothing and reported zero misses when there was one.
Part 8: Models, Latency and Keeping a Fleet Consistent
Lesson 33: Model choice mattered far less than we assumed
We benchmarked four language models on an identical prompt and suite. Scores ranged from 85% to 96%, which looks decisive. Then we analysed 79 real calls across 16 configuration versions and found the models practically indistinguishable, with defect scores of 3.3 versus 3.1. Every recurring defect was prompt-level or config-level, not model-level.
A later head-to-head did show one genuine difference, and it was about consistency, not score:
Stability is the thing worth measuring, and a single run cannot see it. We also blamed one model for an "all attempts exhausted" infrastructure failure at about a 2% rate, then found its replacement did the same at about 1.5%. It was platform noise, and we had to retract the verdict. Benchmark before tuning the prompt, not after, and check the per-minute price before blaming the prompt for cost.
Lesson 34: Latency is usually not where you think
On one line, 39% of the 3 to 4 second pauses were in none of the obvious places: not turn detection, not the language model, not speech synthesis. They were silent webhook calls, where the agent fired a tool and said nothing while it waited.
- Measure the gap the caller actually experiences, then decompose it.
- Model cold starts vary wildly: 4.07 seconds and then 0.79 seconds on the same model. Measure across many calls before optimising anything.
- A long dead-air gap followed by a change in behaviour mid-call is usually a failover timeout, not slowness. Ours was set to 8 seconds on one agent, which produced exactly the "huge delays" complaints we were chasing.
- A failover chain whose first entry is the model you are already running is a no-op. Check it.
Lesson 35: The same fix does not propagate across a fleet
We keep one agent as the reference, because it carries the most real call volume and therefore finds defects first. Nothing propagates automatically. Every fix on the reference agent is an open gap on every other agent until somebody diffs and ports it.
One audit found a client agent that had received none of the platform-layer work since go-live, months earlier. Two gaps were missing capability, not mis-tuning:
- The end-call action was disabled. The agent could not hang up, so every call ran to the 25-minute maximum unless the caller hung up first.
- The stay-silent action was disabled, which is why none of the noise-handling work could have helped it. A prompt can only ask for behaviour the agent has a tool for.
Add an 8-second failover timeout, no failover chain, an empty filler list with defaults disabled, and no speech-recognition keywords.
The fix is procedural, not technical. Fetch both agents through the API, print the two configs side by side, and close every gap, at go-live and every time the reference agent changes. Do not eyeball it; assert it. Ours is now a baseline table plus a runnable assertion block. And the root cause was our own agent-creation template: its prompt skeleton contained the answer-the-noise line from Lesson 10, and its clone payload dropped the fields that enable end call and stay silent. Anything left out of a clone payload gets nothing, not a sensible default. We had been fixing symptoms on individual agents for months.
Lesson 36: Moving rules out of the prompt is promising, with one sharp edge
The platform added a feature that bundles task-specific instructions into named procedures with triggers, so the agent loads only the relevant one rather than carrying the whole prompt on every turn. We tested it properly: baseline, procedures with a 35% smaller prompt, and procedures with the full prompt kept.
The full prompt with no procedures. The comparison point for the other two.
Within the suite's noise floor, so no measurable cost. But five of eight regressions were the wrong procedure firing, or none at all.
Also within the noise floor, with zero trigger-miss regressions, because the rules still existed in the prompt.
The honest conclusion is inconclusive, because the suite's noise floor (Lesson 5) was bigger than the effect. Two things did come out of it. A 35% prompt reduction cost nothing measurable, which makes it worth pursuing for latency and cost, though not for accuracy. And a trigger miss is catastrophic if you moved the rules out of the prompt. Five of eight regressions in the trimmed version were the wrong procedure firing, or none at all, and because the rule had been deleted from the prompt it no longer existed anywhere. One test regressed because "ask which consultant" had moved into a procedure that never fired.
Treat this class of feature as additive to your prompt, not a migration out of it, and measure trigger routing accuracy separately from behaviour before trusting it with anything.
What a Production AI Voice Agent Costs
The lessons above are where the budget of a voice agent really goes: not the platform fee, but the testing, telephony, payments and monitoring around it. Our engagements start with a $5,000 fixed-scope pilot on one call flow. A single production agent with its integrations typically runs $15,000 to $25,000, and a full system like the one these lessons came from, with its own booking system, payments, forms and team ticketing, sits in the $40,000 to $100,000+ band, as set out in our case study of that system.
For a figure for your own phone lines, our automation quote tool gives a ballpark in about a minute, our pricing page covers the retainers that keep an agent maintained after launch, and our voice AI development service explains how a pilot is scoped.
The Competitor Pulse Check
| Factor | How we operate voice agents | Typical voice AI deployment |
|---|---|---|
| Testing side effects | Mock all write tools; TEST in every human-visible field; one deliberate real send | Mocking left on "selected" or off; realistic test names |
| Measuring a change | 3 to 5 rounds per arm; failures classified before prompt edits | One run, one score, one conclusion |
| Monitoring | Tool-call records and transcripts | Platform success flag and auto-summary |
| Noise and silence | Punctuation-only turns ignored; four-step silence ladder | "Sorry, I didn't catch that" on every knock |
| Numbers and names | Grouped words, said once; spelling is the truth | Repeats on request; trusts the spoken read-back |
| Time | Open or closed computed on the server | "You know the current time" in the prompt |
| Fleet consistency | Every agent diffed against a reference by API | Fixes applied by hand to whichever agent complained |
The Short Version
If you are shipping a voice agent this month:
- Mock every outward-facing side effect, and never trust a mocking mode that lists tools by name.
- Put the word TEST in anything that can reach a human.
- Run 3 to 5 rounds before believing any result.
- Never monitor with the platform's own success flag.
- Punctuation-only turns are noise, not speech, and they must not advance your silence counter.
- Say any number once, as grouped words, and count the digits before reading it back.
- A spoken read-back cannot catch a spelling mistake.
- Never let the model decide what time it is.
- Check every promise in the prompt against the tools actually attached.
- When you add a rule, delete the one it contradicts.
- Diff every agent against your reference agent, at go-live and after every change.
- A successful write is not correct behaviour.
Frequently Asked Questions
Why do AI voice agents fail in production when they passed testing?
Because most tests cannot see the failures that matter. Text simulations cannot test timing, interruptions or a caller reading a number in chunks, single test runs sit inside a noise floor larger than most effects, and the platform's success flag restates what the agent claimed rather than what it did. Real calls, tool-call records and repeated test rounds catch what a passing suite misses.
How should you test an AI voice agent safely?
Mock every tool with an outward-facing side effect, set the mock mode to all rather than selected, put the word TEST in every field a human could see, and send the fewest real messages that prove the behaviour. Then run each test 3 to 5 times before trusting the result, and finish with a small number of real phone calls for anything involving timing.
Which platform do you use to build AI voice agents?
Our medical voice platform, Veda, runs its conversation layer on ElevenLabs Agents, around our own practice management system, doctor schedule overrides, doctor forms and payment system. We chose ElevenLabs for its documented APIs, leading speech models and its testing suite with tool mocking and repeated runs. For other clients we still benchmark platforms, including self-hosted open-source stacks, against their own call recordings.
Does the language model matter much for a voice agent?
Less than most teams assume. Across 79 real calls and 16 configuration versions, the models we compared were practically indistinguishable, and every recurring defect was in the prompt or the configuration. Where models did differ was consistency across repeated runs, which a single benchmark run cannot measure.
How do you stop a voice agent from talking over callers or answering background noise?
Treat a punctuation-only transcript as noise and have the agent stay silent using a stay-silent tool, not a prompt instruction alone. Do not count noise as a pause, and do not count fillers like "mm" as a turn. When a caller is genuinely quiet, escalate through a ladder that starts with silence and never repeats the same sentence.
How do you know whether a voice agent actually did what it told the caller?
Read the tool-call records, not the transcript summary or the platform's success flag. In our message-taking calls, about 1 in 21 confirmed logging to the caller without calling the tool, and the platform marked every one of them successful. Check every promise in the prompt against an attached tool, and use structural gates such as required parameters and backend validation.
How long does it take to get an AI voice agent production-ready?
The build is rarely the long part. The time goes on tuning against real calls, wiring the systems around the agent, and closing the gaps this post describes. Our voice AI hub sets out realistic timelines by use case.
What's Next
If you are building a voice agent yourself, start with the voice AI hub for architecture and use cases, and the call centre orchestration guide for telephony, SIP and transfers. If you want to see the full system these lessons came from, read the Veda case study.
If you would rather not learn these 36 lessons on your own phone lines, our voice AI development service builds and runs voice agents with every one of them already in the baseline.
Talk to us about your voice agent →
We run AI voice agents on live medical phone lines. The findings above come from production calls, carrier logs and post-incident review, not from a lab. No patient data, client identity or contact detail appears in this post.
Muhammad Kashif is co-founder of ValueStreamAI, leading technical delivery and AI strategy. He designs and ships custom agentic AI and healthcare automation systems for clients across the US and UK. More about Muhammad Kashif →
