GenAI System Design
2. Inference, GPUs and Serving

GPU Memory and the KV Cache

What lives in GPU memory during inference, how to calculate KV cache size, and why it, not the weights, usually limits how many users a GPU can serve.

Lesson 2 of 7 12 min

What sits in GPU memory

H100, 80 GB, serving Llama 3.1 8B in BF16Weights 16 GB8B params × 2 bytesKV cache pool ≈ 56 GBeach slot = one 8K-token sequence (≈ 1.07 GB), ≈ 52 fit≈ 8 GBruntimeSame GPU, Llama 3.1 70B in BF16141 GB of weights alone: does not fitOptions: split across 2+ GPUs (tensor parallelism), or quantize to INT4 (≈ 40 GB) to leave room for KV cache.
Whatever the weights leave free becomes KV cache, and KV cache is what limits how many users one GPU can serve.

Three things compete for the 80 GB of an H100:

  1. Weights: parameters × bytes per parameter. This is fixed for a given model and precision.
  2. KV cache: attention keys and values for every token of every active sequence. It grows as conversations get longer and as more users are served at once.
  3. Activations and runtime: temporary buffers for the current step, CUDA context, and the engine's own allocations. Plan for a few GB.

Serving engines like vLLM load the weights and then claim most of the remaining memory (90% by default) as a KV cache pool. How many sequences fit in that pool decides your maximum batch size, and so your throughput.

Why the KV cache exists

Attention lets each new token look at all previous tokens. To do that, the model needs a key and a value vector for every previous token in every layer. Recomputing them for the whole history at every step would make generation quadratic. So the model computes them once, during prefill for prompt tokens and during decode for each new token, and keeps them. That stored state is the KV cache.

Calculating it

KV bytes per token = 2 × layers × KV heads × head dim × bytes per value

The 2 is for K and V. The inputs come straight from the model's config.json:

ModelLayersKV headsHead dimPer token (FP16)8K-token sequence
Llama 3.1 8B328128128 KB1.07 GB
Llama 3.1 70B808128320 KB2.7 GB
Llama 3.1 405B1268128504 KB4.2 GB

Then:

Max concurrent sequences ≈ (usable GPU memory − weights − overhead) ÷ (KV per token × tokens per sequence)

For Llama 3.1 8B on one H100: (≈ 76 − 16 − 4) GB ÷ 1.07 GB ≈ 52 sequences at 8K tokens, or about 13 at 32K tokens.

KV cache calculator

See how context length and concurrent users eat GPU memory, and how many sequences one GPU can hold.

Weights 15.0 GBKV cache 16.0 GBActivations + runtime 4.65 GBDashed line: usable GPU memory (76.0 GB)
KV cache per token
128 KB
2 × layers × kv_heads × head_dim × bytes
KV cache per sequence
1.00 GB
at 8,192 tokens
Max concurrent sequences
59
at this context length
Memory needed now
35.6 GB
fits

Techniques that shrink the KV cache

Grouped-query attention (GQA). Older models had one K/V head per attention head, 64 for a 70B model. GQA shares each K/V head across a group of query heads. Llama 3's 8 KV heads instead of 64 cut the KV cache 8× with little quality loss. It is why modern models can offer long contexts at all. Some models go further with multi-head latent attention (MLA, used by DeepSeek), which compresses K and V into a smaller latent vector.

KV cache quantization. Storing K and V in FP8 instead of FP16 halves the cache, which doubles the concurrent sequences or the context length. Quality impact is usually small, and most engines support it as a flag.

Paged allocation. Instead of reserving a max-length buffer per request, the cache is split into small blocks allocated on demand. The inference engines lesson covers PagedAttention.

Prefix sharing. Requests that share a prefix, such as a long system prompt or the same document, can point at the same KV blocks instead of storing copies.

Sliding-window and hybrid attention. Some models attend only to recent tokens in most layers, which caps KV growth for those layers.

When the KV cache runs out

If new requests arrive and the pool is full, the engine must choose between:

  • Queueing new requests, which raises TTFT.
  • Preempting running sequences: evict their KV cache and recompute it later, or swap it to CPU memory. This raises their latency.
  • Rejecting requests with a 429 or 503 so the gateway can route elsewhere.

Key takeaways

  • GPU memory holds weights (fixed), KV cache (grows with tokens × concurrent sequences), and a few GB of activations and runtime buffers.
  • KV cache per token = 2 × layers × KV heads × head dim × bytes per value.
  • Grouped-query attention, KV quantization and paged allocation all exist to fit more KV cache per GPU.
  • Maximum concurrency ≈ (GPU memory − weights − overhead) ÷ KV cache per sequence.

Go deeper

Finished reading? Mark it done to track your progress.