AI Engineering

Building Production-Ready Voice AI Agents on AWS: From Prototype to Production

Voice AI AWS AI Agents

I gave this talk at AWS Community Day DMV 2026 at Amazon HQ2 on October 2, 2026. This is the recap.

Wanjohi Christopher presenting the slide “Two ways to build a voice agent” at AWS Community Day DMV 2026, Amazon HQ2

Presenting “Two ways to build a voice agent”, cascade vs speech-to-speech, at Amazon HQ2.

The talk opened with one sentence: these lessons came from an agent that took real phone calls. I rebuilt it on AWS for the talk, one WebSocket, one Fargate task, and measured everything instead of estimating it.

The scenario

Airline rebooking. One caller, one recording, airport ambience in the background. The caller says three things:

  • “My flight to Baltimore just got cancelled. The six-forty.”
  • “Confirmation K as in kilo, seven, R, two, nine, X.”
  • “And I had a checked bag on the original.”

Four slots to fill. One of them is hard.

Watch it fail

Same audio, same commit, three runs, four booleans. Run A is the naive build: it cuts the caller off, guesses, and answers itself. Run B is the shipped build: fixed, but it cannot be interrupted. Run D is the 2026 build: correct, and 736 ms faster.

Three runs is the whole method. If you cannot replay the same audio against different builds and compare the numbers, you are guessing.

The confidence gate: ask, or guess?

On one take, Transcribe returned “…confirmation is K 7 or 29 X.” with a confidence of 0.33. With the gate off, the agent writes K7OR29X into the booking. No flag, no warning. The wrong passenger flies.

With the gate on at 0.75, the agent refuses to write it. It records the slot as flagged and asks the caller: “K-7-O-R-2-9-X, correct?”

The gate costs one extra turn every time it fires. That is the trade, stated plainly: one turn against a wrong booking reference.

The echo guard: the fix that shipped a walkie-talkie

The caller said five things. Run B heard two:

  1. “…my flight to Baltimore just got cancelled.” (heard)
  2. “I need to get there tonight, or tomorrow morning.” (deaf)
  3. “No, wait, tonight, if there’s anything at all.” (deaf)
  4. “Confirmation K as in kilo, seven, R, two, nine, X.” (heard)
  5. “And I had a checked bag on the original.” (deaf)

The echo guard stopped the agent from answering itself. It also stopped the caller from being heard for 7.3 seconds a turn. The fix shipped a walkie-talkie: one side talks, the other side waits.

Where the seconds actually go

One turn, measured end to end:

  • Endpointer holds: 1500 ms
  • Transcribe final: 17 ms
  • Bedrock first token: 482 ms
  • Lambda lookup: 1851 ms
  • Filler audio out: “Let me check…”
  • Answer synthesis: 219 ms
  • Caller is deaf: 7.3 s unheard

The caller hears something at 1088 ms. The lookup finishes 1.6 s later. That filler is what keeps the call from feeling dead while the lookup runs.

Every fix costs something

Every production fix, with its price tag:

FixWhat it fixedWhat it costs
Endpointer 700 to 1500 msStopped cutting the caller off+800 ms, every utterance
Confidence gateStopped a wrong booking reference+1 turn, when it fires
Echo guardStopped the agent answering itself7250 ms, cannot interrupt
Filler at tool decisionCaller hears something in 1088 ms-1620 ms, a fix that pays
Semantic endpointingCorrect and fast vs the threshold-736 ms

Three of those fixes cost you time, a turn, or the ability to interrupt. Two give time back. The point is not which direction, it is that you know the number before you ship.

Instrument what you cannot see

Every number in the talk has a CloudWatch metric name. The rig emits them with CloudWatch EMF straight from a Fargate task: metrics from stdout, no PutMetricData calls.

The rule on that slide: if you cannot see it, you cannot defend it. When someone asks why the agent feels slow, you want a metric name, not an opinion.

Cascade vs speech-to-speech

About 90% of voice agents still ship a cascade: Transcribe to Bedrock to Polly, 482 ms here, 219 ms there. Every seam is a place to measure, and a place to pay.

Nova 2 Sonic is the other option: one bidirectional stream, faster. The 2026 answer is hybrid. Speech-to-speech where latency dominates. Cascade where you need tools.

Three takeaways

  1. Every fix costs something. The gate adds a turn. The guard costs the ability to interrupt. Ship the fix with the price tag attached.
  2. Measure the seam, not the guess. Everyone blames text-to-speech. 59% was the model. Find the actual seam before you optimize.
  3. Read the words, not the silence. A pause mid-sentence and a pause at the end are identical to the endpointer. Your system has to tell them apart.

Dig deeper

From the talk’s resource list, worth your time next: Amazon Nova 2 Sonic on Bedrock (bidirectional streaming), Transcribe item confidence (what the clarification beat needs), and CloudWatch EMF (metrics from stdout, no PutMetricData). The rig itself is Transcribe to Bedrock to Polly on Fargate, with a measurement harness that reproduces every number, and my notes on what broke, including the fix that only looked fixed.