Blog · · 5 min read

Voice AI Latency: Where the Time Goes and How to Cut It

Voice AI latency is the sum of turn detection, STT, the LLM, TTS and transport. How to measure each stage, and the fixes that matter most, starting with streaming.

Key points

  • Measure voice AI latency from the moment the caller stops speaking to the moment the first audio of the reply goes out, and split it into turn detection, speech-to-text, LLM, text-to-speech and transport.
  • The biggest win is streaming between stages instead of waiting for each to finish. Send the LLM output to speech synthesis one sentence at a time and keep the first sentence short.
  • You can't remove the wait on external API calls, so fill it with a short spoken line such as "Let me check that for you" instead of silence.

A voice agent feels slow because five stages sit between the end of the caller's sentence and the first sound of the reply: turn detection, speech-to-text, the LLM, text-to-speech and transport. To cut latency, measure each stage separately and fix the longest one first. In most pipelines, the changes that matter are streaming between stages rather than waiting for each to finish, and keeping the first sentence of every reply short.

The bar set by human conversation

People answer each other remarkably fast. A study of conversations in ten languages found the mean gap before a response to a yes/no question was about 208 ms across the dataset (and as low as 7 ms for Japanese) (Stivers et al., PNAS, 2009). Humans start planning their reply before the other person finishes.

A pipeline agent can't do that: it has to be fairly sure the caller has finished before it starts. So the realistic goal isn't human parity. It's getting some audio back before the caller wonders whether anyone is there.

The five stages

Latency is the sum of these stages, on the phone or on the web:

Stage What happens Why it gets long
1. Turn detection Silence is detected and the turn is judged complete Long silence thresholds, padding to avoid cutting people off
2. Speech-to-text final The last words are transcribed and finalized Recognition starts only after the utterance ends
3. LLM first token The model reads the prompt and history and starts replying Long prompts, long history, large models, tool calls
4. TTS first audio The reply is turned into audio Synthesis waits for the full reply
5. Transport Audio travels to the caller Distance between servers and caller, phone network, audio buffers

Stage 1 is a trade-off, not just a speed knob: shorten it too much and the agent jumps in whenever someone pauses mid-sentence. That's a turn-taking problem, covered in Barge-in and turn-taking.

Measuring each stage

Before changing anything, timestamp these events on every turn:

  1. Caller audio stopped (VAD detects silence)
  2. Turn judged complete
  3. Final transcript available
  4. First LLM token received
  5. First TTS audio chunk received
  6. First audio chunk sent to the caller

The differences are your stage durations. Look at the slow tail (p90 or p95), not just the average. It's the occasional long pause that makes a call feel broken.

python
# Turn latency breakdown from event timestamps (seconds)
STAGES = [
    ("turn_end", "speech_stopped", "turn_detected"),
    ("stt", "turn_detected", "transcript_final"),
    ("llm_first_token", "transcript_final", "llm_first_token"),
    ("tts_first_audio", "llm_first_token", "tts_first_audio"),
    ("send", "tts_first_audio", "audio_sent"),
]

def breakdown(ts: dict[str, float]) -> dict[str, int]:
    out = {name: round((ts[end] - ts[start]) * 1000) for name, start, end in STAGES}
    out["total"] = round((ts["audio_sent"] - ts["speech_stopped"]) * 1000)
    return out  # milliseconds

How to cut it

Roughly in order of impact.

Stream between stages

The single biggest improvement is not waiting for the previous stage to finish:

  • Use streaming speech recognition, so transcription happens while the caller is still talking.
  • Consume LLM output token by token rather than waiting for the full reply.
  • Split the output at sentence boundaries and send each sentence to TTS as soon as it's complete.

If you synthesize only after the full reply exists, every extra sentence delays the first sound. Sentence-by-sentence synthesis makes time-to-first-audio mostly independent of reply length.

Keep the first sentence short

With sentence streaming, the first sentence determines when the caller hears anything. Instruct the agent to open with a short acknowledgment before the substance. It also sounds more natural.

  • Slow: "Of course, I can help you reschedule your appointment, so could you please tell me the date of your current booking and the new date you'd prefer?"
  • Fast: "Sure." "What's the date of your current booking?"

Give the LLM less to read

Time to first token grows with the amount of input.

  • Strip instructions that don't apply to this use case.
  • Write FAQ answers as short, direct sentences.
  • On long calls, summarize older turns to keep the history compact.

Fill tool waits with words

When the agent calls an external API, such as an availability lookup, you can't make that wait disappear. Silence on a phone line sounds like a dropped call. Have the agent say "Let me check that for you" before calling, cap the wait, and if the cap is hit, apologize and offer a callback. More in Function calling in live calls.

Stop instantly on barge-in

This one is the reverse problem, but it feels the same to the caller: if they start talking and the agent keeps going, the conversation drags. On the phone, audio you've already sent is buffered downstream. With Twilio Media Streams, sending a clear message empties that buffer (WebSocket messages).

Put processing close to the call

Every round trip between your media server and the STT, LLM and TTS APIs adds time. Run the media server in a region close to both the APIs you use and the telephony edge your calls come through.

Doing this with voicast

With voicast, the platform runs the pipeline: turn detection, speech recognition, the LLM and speech synthesis. When the caller starts talking, the agent stops and listens (barge-in). For in-call tools, it says it's checking, then waits up to 8 seconds for your API; on a timeout or error, it apologizes and arranges a callback from staff. Your part is keeping the conversation setup's instructions and FAQ answers tight, and making your tool endpoints respond quickly.

FAQ

Where should I start measuring voice AI latency?

From the moment the caller's audio stops to the moment the first audio of the reply is sent. Then split that into turn detection, speech-to-text, LLM, text-to-speech and send, and work on the longest stage first.

Will a shorter end-of-turn timeout make the agent faster?

Yes, but the agent will start talking over people who pause mid-thought or read out a phone number in chunks. Tune it per use case and pair it with fast barge-in handling.

Is switching to a faster LLM the best fix?

Measure first. If the LLM stage dominates, it helps. But if you still wait for the full reply before synthesizing speech, long replies will stay slow. Streaming between stages usually comes first.

Related posts