GenAI System Design
2. Inference, GPUs and Serving

Why LLM Inference Is Different

Prefill vs decode, why generation is limited by memory bandwidth rather than compute, and what TTFT and TPOT mean for system design.

Lesson 1 of 7 10 min

A request is two very different workloads

When a request reaches a model server, the GPU does two phases of work:

One request on the GPUPrefill2,000 prompt tokens in parallelcompute-boundDecode: 1 token per step, memory-bandwidth-boundTPOT ≈ 10–30 mstimeTime to first token (TTFT)
Prefill processes the whole prompt in one pass. Decode then produces one token per step, re-reading every weight each time.

Prefill. The model processes every prompt token at once and builds the attention state (the KV cache) for each of them. Thousands of tokens flow through big matrix multiplications in parallel, which keeps the GPU's arithmetic units busy. Prefill is compute-bound. Its duration grows with prompt length and determines time to first token (TTFT).

Decode. The model generates the answer one token at a time. Each new token depends on the previous one, so steps cannot run in parallel within a sequence. Every step pushes a single token per sequence through the entire network. Decode is memory-bandwidth-bound. Its step time is the time per output token (TPOT), also called inter-token latency.

Why decode is memory-bound

To produce one token, the GPU must multiply that token's activations by every weight matrix in the model. So each decode step reads all the weights from GPU memory.

Compare the two resources on an H100 for Llama 3.1 8B in BF16:

  • Memory: 16 GB of weights ÷ 3.35 TB/s ≈ 4.8 ms just to stream the weights once.
  • Compute: ≈ 2 FLOPs per parameter per token = 16 GFLOPs ÷ ≈ 990 TFLOPS ≈ 0.016 ms.

At batch size 1 the arithmetic units sit idle about 99% of the time, waiting for bytes. The ratio of FLOPs to bytes moved is called arithmetic intensity. At batch 1, decode has an intensity of about 1 FLOP per byte, while an H100 needs about 300 to be compute-bound.

Batching: the free lunch

The weights are read once per step, regardless of how many sequences are in the batch. If 32 users decode together, the GPU reads the 16 GB once and does 32 tokens of arithmetic with it. Step time barely moves, but you produce 32× the tokens.

That is why every serving engine batches aggressively, and why throughput per GPU and latency per user are separate numbers:

  • At batch 1 (Llama 3.1 8B, H100): ≈ 150 tokens/s for one user and 150 tokens/s total.
  • At batch 64 (2K-token contexts): ≈ 70 tokens/s per user, but ≈ 4,500 tokens/s total.

The limits come from two places. Each sequence's KV cache must also be read every step, so long contexts add bytes. Eventually compute catches up with bandwidth. The batching lesson lets you explore both.

TTFT and TPOT as SLOs

Production serving targets are usually stated at a percentile, for example:

MetricTypical interactive targetDriven by
TTFT p95< 500 ms – 1 sQueueing + prefill (prompt length, model size)
TPOT p95< 30–50 ms (≥ 20–30 tokens/s)Batch size, model size, memory bandwidth
Throughputas high as possibleBatch size, GPU memory for KV cache

These pull against each other. Bigger batches raise throughput but slow every user's TPOT. A long prefill that joins a running batch can stall everyone's decode for a step, which shows up as a latency spike. Engines mitigate this with chunked prefill, which splits long prompts into pieces interleaved with decode steps. Large deployments go further with disaggregated serving, putting prefill and decode on separate GPU pools.

What this means for design

  • Long prompts cost TTFT and money. Long outputs cost total latency. Put limits on both.
  • Self-hosted capacity is measured in tokens per second per GPU at a target TPOT, not in QPS.
  • GPU memory is the scarce resource. Anything that shrinks weights or the KV cache lets you batch more users per GPU.

Key takeaways

  • Generation has two phases. Prefill processes the prompt in parallel and is compute-bound. Decode produces one token per step and is memory-bandwidth-bound.
  • Prefill sets time to first token (TTFT). Decode sets time per output token (TPOT).
  • Each decode step reads every weight in the model, so at batch size 1 a GPU is almost entirely waiting on memory.
  • Batching shares each weight read across many users. This is the main lever for throughput.

Go deeper

Finished reading? Mark it done to track your progress.