Prefill vs Decode: TTFT and TPOT
The two phases of every LLM call seen from the application side, how to measure and budget TTFT and TPOT, and why streaming changes UX and architecture.
Two phases, two latencies
Every LLM call, hosted or self-hosted, has the same shape:
- Prefill processes the whole prompt in one parallel pass. Its cost grows with prompt length. It ends when the first output token is ready, so it sets time to first token (TTFT).
- Decode generates output tokens one at a time, each depending on the last. The gap between tokens is the time per output token (TPOT), also called inter-token latency.
Total latency ≈ TTFT + (output tokens − 1) × TPOT
Module 2 explains the hardware reason: prefill is compute-bound and decode is memory-bandwidth-bound. This lesson is about what those numbers mean for the application.
Typical numbers
Where latency comes from, and what fixes it
| If this is slow | Likely cause | Fixes |
|---|---|---|
| TTFT | Long prompt, queueing at the provider, cold cache | Shorter prompts, prompt caching, a smaller model, provisioned throughput, fewer sequential steps before the call |
| TPOT | Big model, busy server, long context | A smaller or faster model, a different provider or region, speculative decoding (self-hosted) |
| Total | Long outputs, many sequential calls | Cap output length, ask for concise formats, run calls in parallel, stream |
Streaming
With streaming, the server sends tokens as they are generated, usually over server-sent events (SSE), instead of waiting for the full response.
- For humans, perceived latency drops from total time to TTFT. People read at about 5–8 tokens/s, so any decode speed above about 20 tokens/s feels instant once text is flowing.
- For machines, streaming doesn't help if you need the complete response, for example to parse JSON or call a tool. Latency is still total time. This is why agents feel slow: every step pays the full generation time.
Streaming also changes your architecture:
- Connections are long-lived. A 20-second response means a 20-second HTTP connection. Load balancers, proxies and serverless platforms need timeouts and buffering configured for this.
- Errors can happen mid-stream. The client must handle a stream that stops halfway, and your API should signal that with an error event.
- Guardrails get harder. You can't moderate a full answer before showing it. Common approaches are to check chunks as they stream (with a small delay buffer), or to hold back specific risky outputs.
- Billing and logging must be computed when the stream completes or is cancelled. If the user closes the tab, cancel the upstream request so you stop paying for tokens nobody reads.
Measuring latency properly
Log per request: model, input tokens, output tokens, cached tokens, TTFT, total time, and TPOT derived as (total − TTFT) ÷ (output tokens − 1). Then:
- Track p50 and p95 separately for TTFT and TPOT. Averages hide the tail.
- Bucket by prompt size. A TTFT regression that only affects 50K-token prompts disappears in a global average.
- Alert on TTFT p95 for interactive features, and on total time for agent and batch features.
Key takeaways
- An LLM call has two phases. Prefill reads the prompt and sets time to first token (TTFT). Decode writes the answer one token at a time and sets time per output token (TPOT).
- End-to-end latency ≈ TTFT + output tokens × TPOT. Output length is usually the biggest term.
- Streaming shows tokens as they are generated. It hides decode time for humans, but not for machines consuming the full response.
- Measure TTFT, TPOT and total time as separate percentiles, per model and per prompt size.