Quantization and Speculative Decoding
Two ways to generate tokens faster and cheaper without buying more GPUs, what each costs in quality, and when to use them.
Quantization: fewer bytes per weight
Models are trained in 16-bit formats (BF16). Quantization stores the weights in 8, 4 or even fewer bits, with scale factors that map them back to the right range.
What quantization buys you
Fewer bytes per weight means fewer GPUs and faster decode, because decode is limited by how fast weights stream out of memory.
| Precision | Weights | GPUs | Single-user speed | Quality |
|---|---|---|---|---|
| BF16 | 141 GB | 4 | 66 tok/s | Reference quality. |
| FP8 | 71 GB | 2 | 66 tok/s | Near-lossless on H100/B200-class GPUs; the default for production serving. |
| INT4 (AWQ/GPTQ) | 37 GB | 1 | 62 tok/s | Small drops on hard reasoning and code; usually fine for chat and RAG. |
| 3-bit | 35 GB | 1 | 66 tok/s | Noticeable degradation. Local/edge use, not production APIs. |
Speeds are roofline estimates at batch size 1. Real INT4 kernels gain less than the byte ratio suggests because of dequantization overhead, especially at large batch sizes where decode becomes compute-bound.
It helps in three ways at once:
- Less memory: a 70B model drops from 140 GB (BF16) to 70 GB (FP8) or ≈ 38 GB (INT4). You need fewer GPUs per replica.
- Faster decode: decode is memory-bound, so half the bytes means close to twice the tokens/s at small batch sizes.
- Bigger batches: freed memory becomes KV cache, which means more concurrent users per GPU.
Formats you should know
| Format | Bits | Typical use | Quality |
|---|---|---|---|
| BF16 / FP16 | 16 | Training, reference serving | Baseline |
| FP8 (E4M3) | 8 | Production serving on H100/H200/B200, with native hardware support | Near-lossless for most models |
| INT8 (weights, sometimes activations) | 8 | Older GPUs (A100), CPUs | Near-lossless |
| INT4: AWQ, GPTQ | ≈ 4.25 | Fitting big models on fewer or smaller GPUs | Small drops, most visible in maths, code and long reasoning |
| FP4 (NVFP4, MXFP4) | ≈ 4.25 | Blackwell-class GPUs, some models shipped natively in it | Close to FP8 when the model is trained or calibrated for it |
| GGUF k-quants (Q4_K_M, …) | 3–8 | llama.cpp / local inference | Varies by level |
Weight-only vs weight-and-activation
Weight-only schemes (AWQ, GPTQ) store weights in 4 bits but compute in 16 bits after unpacking. That saves memory and bandwidth, but not compute, so gains shrink at large batch sizes. Weight + activation schemes (FP8, FP4) run the matrix multiplications in low precision on hardware that supports it. That also speeds up compute-bound prefill.
KV cache quantization
The KV cache can also be stored in FP8, which halves it. For long-context workloads, where KV dominates memory, this often helps more than quantizing the weights.
Speculative decoding
At small batch sizes, each decode step leaves the GPU's compute almost idle while weights stream from memory. Speculative decoding puts that idle compute to work:
- A cheap draft proposes the next k tokens, for example 4.
- The big target model checks all k in one forward pass. This is like a mini prefill, and costs about the same memory traffic as generating one token.
- The target accepts the longest prefix of draft tokens that matches what it would have sampled, then adds one token of its own.
If on average 3 of 4 drafts are accepted, you get about 4 tokens per expensive step instead of 1. A rejection-sampling rule guarantees the output distribution is identical to the target model alone. Quality is not traded away. Only speed changes.
Where drafts come from
- A small model of the same family, for example an 8B drafting for a 70B.
- Extra prediction heads trained on the target (Medusa, EAGLE). No separate model to serve, and high acceptance rates.
- N-gram or prompt lookup, which copies spans from the prompt. This works very well for editing, code refactors and RAG answers that quote sources.
When it helps and when it doesn't
- Typical speedups are 1.5–3× on TPOT at low to moderate batch sizes.
- It helps less at high batch sizes. The GPU is no longer idle, and verification work for rejected tokens costs real throughput.
- Acceptance depends on how predictable the text is. Code and structured output are more predictable than creative writing.
Key takeaways
- Quantization stores weights (and optionally the KV cache) in fewer bits. FP8 is near-lossless and is the production default on modern GPUs. INT4 trades some quality for 4× smaller weights.
- Fewer bytes per weight speeds up memory-bound decode and frees memory for bigger batches.
- Speculative decoding has a cheap draft propose several tokens that the big model verifies in one step. The output distribution is unchanged.
- Speculative decoding helps most at low batch sizes, where the GPU has spare compute.