Batching and Throughput
How batching trades per-user speed for total throughput, why continuous batching beats static batching, and how to pick a batch size for a latency target.
Throughput vs latency, one GPU at a time
The previous lesson showed that a decode step costs about the same whether it serves 1 sequence or 32. Here is how that plays out as the batch grows:
Batching simulator
One decode step reads every weight once, whether it serves 1 sequence or 64. Batching shares that read, until KV cache reads or compute take over.
Throughput climbs almost linearly at small batches while per-user speed barely drops: that is free capacity. Past the knee, KV cache reads (long contexts) or compute (short contexts, big batches) dominate, and every extra user slows everyone down.
Watch three regions:
- Small batches (1–16): total throughput rises almost linearly and per-user speed barely drops. This capacity is essentially free, so never run a production GPU at batch 1.
- The knee: each sequence's KV cache must also be read every step. With long contexts, KV reads start to rival the weight reads, and step time grows with batch size.
- The ceiling: either the KV cache pool is full (red line), or the batch is big enough that compute, not bandwidth, becomes the limit. More sequences then just slow everyone down.
Static batching and why it wastes GPUs
The classic approach, borrowed from training, gathers N requests, runs them together until all finish, then takes the next N. Output lengths vary enormously ("yes" vs a 2,000-word essay), so most slots sit idle waiting for the longest sequence, and queued requests wait even longer.
Static vs continuous batching
Eight requests with different answer lengths share a GPU that decodes 4 sequences per step. Each cell is one decode step for one slot.
With static batching, short requests (R2) finish and leave their slot idle until the longest request (R3) is done, and R5 to R8 wait in the queue. Continuous batching, also called in-flight batching, admits a new request the moment a slot frees up. vLLM, SGLang, TGI and TensorRT-LLM all do this.
Continuous batching (also called in-flight or iteration-level batching) makes scheduling decisions every decode step. When a sequence finishes, a waiting request takes its slot on the next step. The 2022 Orca paper introduced it, and every modern engine does it. vLLM's authors reported throughput gains of several times to over 20× over naive static batching, depending on the workload.
How a scheduler thinks
Each step, the engine's scheduler decides which sequences run, subject to:
- A KV cache budget. A new request is admitted only if there are free blocks for its prompt, plus some room to grow.
- A token budget per step. Many engines cap the total tokens processed per step, for example 8K. A long prompt is then split into chunks (chunked prefill) instead of stalling every decoding sequence for one huge prefill step.
- Priorities and fairness. Interactive traffic before batch jobs, and per-tenant limits.
If memory runs out mid-generation, the scheduler preempts a sequence. It drops the sequence's KV cache and recomputes it later, or swaps it to CPU memory.
Choosing a batch size
Work backwards from the product requirement:
- Set a TPOT target from the UX. For example, ≥ 30 tokens/s per user means TPOT ≤ 33 ms.
- Find the largest batch that meets it. Use a benchmark or the roofline estimate: step time ≈ (weight bytes + batch × KV bytes per sequence) ÷ effective bandwidth.
- Check memory: batch × KV per sequence must fit in the pool at your typical context length.
- Throughput per GPU = batch ÷ step time. This is the number you divide your peak token rate by in capacity planning.
Separate pools for separate goals
At scale, teams often run the same model in two configurations:
- Latency pool: smaller batches, more replicas, speculative decoding, for chat and agents.
- Throughput pool: maximal batches, for summarisation backfills, evals, embeddings and batch-API traffic.
The gateway routes by request type or priority. This is a strong point to raise when designing a serving platform.
Key takeaways
- Batching shares each weight read across many sequences. Throughput climbs almost linearly until KV cache reads or compute take over.
- Static batching wastes GPU time waiting for the longest sequence. Continuous batching refills free slots at every step.
- Choose batch size from a TPOT target, then check that the KV cache for that batch fits in memory.
- Throughput-optimised and latency-optimised deployments of the same model can differ several-fold in cost per token.