GPU Memory and the KV Cache
What lives in GPU memory during inference, how to calculate KV cache size, and why it, not the weights, usually limits how many users a GPU can serve.
What sits in GPU memory
Three things compete for the 80 GB of an H100:
- Weights: parameters × bytes per parameter. This is fixed for a given model and precision.
- KV cache: attention keys and values for every token of every active sequence. It grows as conversations get longer and as more users are served at once.
- Activations and runtime: temporary buffers for the current step, CUDA context, and the engine's own allocations. Plan for a few GB.
Serving engines like vLLM load the weights and then claim most of the remaining memory (90% by default) as a KV cache pool. How many sequences fit in that pool decides your maximum batch size, and so your throughput.
Why the KV cache exists
Attention lets each new token look at all previous tokens. To do that, the model needs a key and a value vector for every previous token in every layer. Recomputing them for the whole history at every step would make generation quadratic. So the model computes them once, during prefill for prompt tokens and during decode for each new token, and keeps them. That stored state is the KV cache.
Calculating it
KV bytes per token = 2 × layers × KV heads × head dim × bytes per value
The 2 is for K and V. The inputs come straight from the model's config.json:
| Model | Layers | KV heads | Head dim | Per token (FP16) | 8K-token sequence |
|---|---|---|---|---|---|
| Llama 3.1 8B | 32 | 8 | 128 | 128 KB | 1.07 GB |
| Llama 3.1 70B | 80 | 8 | 128 | 320 KB | 2.7 GB |
| Llama 3.1 405B | 126 | 8 | 128 | 504 KB | 4.2 GB |
Then:
Max concurrent sequences ≈ (usable GPU memory − weights − overhead) ÷ (KV per token × tokens per sequence)
For Llama 3.1 8B on one H100: (≈ 76 − 16 − 4) GB ÷ 1.07 GB ≈ 52 sequences at 8K tokens, or about 13 at 32K tokens.
KV cache calculator
See how context length and concurrent users eat GPU memory, and how many sequences one GPU can hold.
Techniques that shrink the KV cache
Grouped-query attention (GQA). Older models had one K/V head per attention head, 64 for a 70B model. GQA shares each K/V head across a group of query heads. Llama 3's 8 KV heads instead of 64 cut the KV cache 8× with little quality loss. It is why modern models can offer long contexts at all. Some models go further with multi-head latent attention (MLA, used by DeepSeek), which compresses K and V into a smaller latent vector.
KV cache quantization. Storing K and V in FP8 instead of FP16 halves the cache, which doubles the concurrent sequences or the context length. Quality impact is usually small, and most engines support it as a flag.
Paged allocation. Instead of reserving a max-length buffer per request, the cache is split into small blocks allocated on demand. The inference engines lesson covers PagedAttention.
Prefix sharing. Requests that share a prefix, such as a long system prompt or the same document, can point at the same KV blocks instead of storing copies.
Sliding-window and hybrid attention. Some models attend only to recent tokens in most layers, which caps KV growth for those layers.
When the KV cache runs out
If new requests arrive and the pool is full, the engine must choose between:
- Queueing new requests, which raises TTFT.
- Preempting running sequences: evict their KV cache and recompute it later, or swap it to CPU memory. This raises their latency.
- Rejecting requests with a 429 or 503 so the gateway can route elsewhere.
Key takeaways
- GPU memory holds weights (fixed), KV cache (grows with tokens × concurrent sequences), and a few GB of activations and runtime buffers.
- KV cache per token = 2 × layers × KV heads × head dim × bytes per value.
- Grouped-query attention, KV quantization and paged allocation all exist to fit more KV cache per GPU.
- Maximum concurrency ≈ (GPU memory − weights − overhead) ÷ KV cache per sequence.