GenAI System Design
10. GenAI System Design Problems

Designing a Real-Time Voice Agent

A full walkthrough for a phone and voice assistant that responds in under a second, covering the streaming ASR-LLM-TTS pipeline vs speech-to-speech models, the latency budget, turn detection and barge-in, telephony, scaling per concurrent call, and evaluation.

Lesson 9 of 10 14 min

1. Requirements

Functional

  • Answer inbound phone calls and web voice sessions for customer support
  • Understand natural speech, answer questions (RAG), and take actions (tools)
  • Handle interruptions, and transfer to a human with context
  • Multiple languages (start with English and Spanish)

Non-functional

  • Voice-to-voice latency (the user stops speaking to the agent starting to speak) under 1 s at p50 and 1.5 s at p95
  • 10K concurrent calls at peak
  • Call recording and transcripts per compliance rules, and PII handling (card numbers)

2. Operating point

Latency is the product. Beyond about 1.5 s, callers talk over the agent and the experience breaks. Use a fast model, and accept slightly lower reasoning quality, with escalation to a human for complex cases.

3. The latency budget

StageBudget
Network and telephony jitter buffer50–100 ms
End-of-turn detection (silence plus semantic cue)200–300 ms
ASR final transcript (streaming, mostly done already)50–100 ms
LLM time to first token (small or fast model, cached prefix)200–350 ms
TTS time to first audio (streaming)100–200 ms
Total≈ 600–1,050 ms

To stay inside this:

  • Stream everything: ASR partials, LLM tokens into TTS sentence by sentence, TTS audio chunks.
  • Start the LLM early on a confident partial transcript (speculatively), and cancel if the user keeps talking.
  • Short first sentence: prompt the model to lead with a brief acknowledgement, so TTS can start.
  • Filler for slow tools: "Let me check that order for you…" while the tool call runs.
  • Co-locate ASR, LLM and TTS in the same region as the media servers.

4. Architecture

Caller
PSTN / SIP / WebRTC
Media server
audio streams, recording
Streaming ASR
partials + finals
Turn manager
VAD, endpointing, barge-in
LLM agent
streaming, tools, RAG
Streaming TTS
audio back to caller
  • Session state service: transcript, agent state and tool results per call, persisted for transfer and compliance.
  • Tools: the same backend tools as the chat support agent (order lookup, returns), with the same identity and limit rules. Caller identity is verified by phone number plus an extra check.
  • Human transfer: SIP transfer with a summary pushed to the agent desktop.

5. Deep dives

Deep dive A: turn-taking and barge-in

  • Endpointing: voice activity detection alone cuts people off mid-thought when they pause. Combine silence duration with a semantic end-of-turn model, a small classifier on the partial transcript ("I want to change my…" is not finished).
  • Barge-in: when the caller speaks while the agent is talking, stop TTS within about 100–200 ms, cancel the in-flight LLM generation, and truncate the conversation history to what was actually spoken, so the model knows what the user heard.
  • Backchannels: ignore "uh-huh" and "okay" as interruptions. Treat "wait, no" as one.
  • Echo cancellation, so the agent's own voice isn't transcribed as user speech.

Deep dive B: cascaded vs speech-to-speech

Cascaded (ASR → LLM → TTS)Speech-to-speech model
LatencySum of stages, ≈ 0.7–1.2 sOften lower
NaturalnessDepends on TTS. Loses the user's toneHears tone and emotion, more natural prosody
Control and debuggabilityText at every stage: easy to log, filter, testHarder to inspect and guard
Tool use, RAG, guardrailsMatureImproving, more limited
Vendor flexibilityMix the best ASR, LLM and TTSTied to one model

A pragmatic choice: cascaded for transactional support (tools, compliance, auditability), and speech-to-speech for conversational experiences where naturalness matters most. Keep the architecture able to swap between them.

Deep dive C: scaling per concurrent call

  • 10K concurrent calls means 10K live ASR streams, 10K TTS streams, and LLM requests at each turn (about every 5–10 s per call, so ≈ 1–2K LLM requests per second at peak).
  • ASR and TTS models run on GPUs with streaming batching. Capacity is measured as concurrent streams per GPU, determined by benchmarking.
  • The LLM needs a low TTFT, so use a small or mid model, prefix caching of the system prompt and conversation, and a latency-optimised pool at moderate batch sizes.
  • Media servers scale horizontally, with sticky sessions per call.

6. Evaluation and compliance

  • Latency p50 and p95 per stage and end to end, measured from real audio timestamps.
  • Task success (resolved without transfer), transfer rate, average handle time, and CSAT from post-call surveys.
  • ASR word error rate on your domain (product names, order numbers), per language and accent.
  • Simulated callers: LLM-driven synthetic callers with scripted goals and speech synthesis, run against every release.
  • PII: pause recording or redact during card capture (hand off to DTMF or a compliant payment IVR), announce recording, and set retention.

7. Bottlenecks and follow-ups

Likely questionAnswer sketch
"Agent talks over people?"Better endpointing (semantic model), faster barge-in cancellation, echo cancellation
"Tool lookup takes 2 s?"Filler speech, prefetch customer data at call start, parallel tool calls
"Accents and noisy lines?"Domain-tuned ASR, keyword boosting for product names, confirm critical details back to the caller
"Cost per minute?"ASR + TTS + LLM tokens + telephony per minute, compared with a human agent's cost per minute

Key takeaways

  • Voice UX needs a response within about 1 second. Every stage must stream, and the budget is split across endpointing, ASR, LLM TTFT and TTS first audio.
  • Turn detection (when the user has finished) and barge-in (the user interrupts) matter as much as model quality.
  • A cascaded pipeline (ASR → LLM → TTS) gives control and tool use. Speech-to-speech models give lower latency and more natural prosody.
  • Capacity is planned per concurrent call, with GPU streams for ASR, LLM and TTS, plus telephony media servers.

Go deeper