GenAI System Design
2. Inference, GPUs and Serving

Quantization and Speculative Decoding

Two ways to generate tokens faster and cheaper without buying more GPUs, what each costs in quality, and when to use them.

Lesson 5 of 7 10 min

Quantization: fewer bytes per weight

Models are trained in 16-bit formats (BF16). Quantization stores the weights in 8, 4 or even fewer bits, with scale factors that map them back to the right range.

What quantization buys you

Fewer bytes per weight means fewer GPUs and faster decode, because decode is limited by how fast weights stream out of memory.

PrecisionWeightsGPUsSingle-user speedQuality
BF16141 GB4
66 tok/s
Reference quality.
FP871 GB2
66 tok/s
Near-lossless on H100/B200-class GPUs; the default for production serving.
INT4 (AWQ/GPTQ)37 GB1
62 tok/s
Small drops on hard reasoning and code; usually fine for chat and RAG.
3-bit35 GB1
66 tok/s
Noticeable degradation. Local/edge use, not production APIs.

Speeds are roofline estimates at batch size 1. Real INT4 kernels gain less than the byte ratio suggests because of dequantization overhead, especially at large batch sizes where decode becomes compute-bound.

It helps in three ways at once:

  1. Less memory: a 70B model drops from 140 GB (BF16) to 70 GB (FP8) or ≈ 38 GB (INT4). You need fewer GPUs per replica.
  2. Faster decode: decode is memory-bound, so half the bytes means close to twice the tokens/s at small batch sizes.
  3. Bigger batches: freed memory becomes KV cache, which means more concurrent users per GPU.

Formats you should know

FormatBitsTypical useQuality
BF16 / FP1616Training, reference servingBaseline
FP8 (E4M3)8Production serving on H100/H200/B200, with native hardware supportNear-lossless for most models
INT8 (weights, sometimes activations)8Older GPUs (A100), CPUsNear-lossless
INT4: AWQ, GPTQ≈ 4.25Fitting big models on fewer or smaller GPUsSmall drops, most visible in maths, code and long reasoning
FP4 (NVFP4, MXFP4)≈ 4.25Blackwell-class GPUs, some models shipped natively in itClose to FP8 when the model is trained or calibrated for it
GGUF k-quants (Q4_K_M, …)3–8llama.cpp / local inferenceVaries by level

Weight-only vs weight-and-activation

Weight-only schemes (AWQ, GPTQ) store weights in 4 bits but compute in 16 bits after unpacking. That saves memory and bandwidth, but not compute, so gains shrink at large batch sizes. Weight + activation schemes (FP8, FP4) run the matrix multiplications in low precision on hardware that supports it. That also speeds up compute-bound prefill.

KV cache quantization

The KV cache can also be stored in FP8, which halves it. For long-context workloads, where KV dominates memory, this often helps more than quantizing the weights.

Speculative decoding

At small batch sizes, each decode step leaves the GPU's compute almost idle while weights stream from memory. Speculative decoding puts that idle compute to work:

  1. A cheap draft proposes the next k tokens, for example 4.
  2. The big target model checks all k in one forward pass. This is like a mini prefill, and costs about the same memory traffic as generating one token.
  3. The target accepts the longest prefix of draft tokens that matches what it would have sampled, then adds one token of its own.

If on average 3 of 4 drafts are accepted, you get about 4 tokens per expensive step instead of 1. A rejection-sampling rule guarantees the output distribution is identical to the target model alone. Quality is not traded away. Only speed changes.

Draft proposes
"the cat sat on"
Target verifies
1 forward pass over 4 tokens
Accept 3
"the cat sat"
+1 from target
"down"
4 tokens produced for the cost of roughly one target step, plus the draft's (much smaller) cost.

Where drafts come from

  • A small model of the same family, for example an 8B drafting for a 70B.
  • Extra prediction heads trained on the target (Medusa, EAGLE). No separate model to serve, and high acceptance rates.
  • N-gram or prompt lookup, which copies spans from the prompt. This works very well for editing, code refactors and RAG answers that quote sources.

When it helps and when it doesn't

  • Typical speedups are 1.5–3× on TPOT at low to moderate batch sizes.
  • It helps less at high batch sizes. The GPU is no longer idle, and verification work for rejected tokens costs real throughput.
  • Acceptance depends on how predictable the text is. Code and structured output are more predictable than creative writing.

Key takeaways

  • Quantization stores weights (and optionally the KV cache) in fewer bits. FP8 is near-lossless and is the production default on modern GPUs. INT4 trades some quality for 4× smaller weights.
  • Fewer bytes per weight speeds up memory-bound decode and frees memory for bigger batches.
  • Speculative decoding has a cheap draft propose several tokens that the big model verifies in one step. The output distribution is unchanged.
  • Speculative decoding helps most at low batch sizes, where the GPU has spare compute.

Go deeper