A framework for testing voice agents before they take live calls, built around a nine-scenario suite covering interruptions, silence, accents, anger, and wrong numbers. Testing scores whole calls against written expected outcomes rather than word-level transcript accuracy, using four measures: task completion, turns to resolution, escalation, and false statements. TTS-based test clips are flagged as insufficient since they lack the timing, disfluencies, and emotional prosody of real callers. The piece also covers prompt regression testing (small edits can swing behavior significantly depending on the model), a three-stage rollout from playground clips to telephony to production telemetry, and a post-launch watchlist for failure modes no suite catches in advance.
Table of contents
Key takeawaysYour test audio is easier than your callersBuild a scenario suite before you build anything elsePass and fail criteria for a conversationTreat the system prompt like codeFrom playground to real phone callsWhat voice agent testing can't catch before launchFAQQuestions this post answers
Why doesn't word error rate (WER) tell you if a voice agent is ready for real calls?
WER only measures transcription accuracy at the word level, so it can't detect whether an agent interrupted a caller, looped on a confirmation, or confidently read back the wrong order number. Voice agent readiness requires scoring the whole conversation against a written expected outcome, not just checking whether the transcript matches a reference. Teams shipping voice agents can track evaluation approaches beyond WER benchmarks on daily.dev.
What scenarios should a voice agent test suite cover before going live?
A nine-scenario suite should include: happy path, caller changing their mind mid-sentence, caller interrupting the agent, silence after a question, out-of-scope requests, background noise with a second voice, strong accents, angry-but-polite callers, and wrong-number callers. Each scenario needs a written expected end state a reviewer can mark without interpretation, such as confirming an appointment exists for the correct time. daily.dev helps engineers building conversational AI compare testing checklists like this one.
Why does editing a system prompt for a voice agent require re-running the entire test suite?
A single word change can flip behavior in scenarios never directly touched, and the same edit can help one model while hurting another. Research on Flan-T5 models found swapping the word excludes for lacks degraded one model by 28% while improving a larger variant by 46%, and 55% of studied API model updates showed no consistent direction across prompts. Testing must use the exact model being deployed. daily.dev surfaces practical guidance for developers managing prompt regressions in production agents.
Share this post