Capacity Planning for Inference
A step-by-step method to size a self-hosted LLM fleet, from traffic to GPUs, with the roofline estimate behind it and a build vs buy comparison.
The method
The roofline estimate behind it
For each decode step, the GPU must:
- Read all weights once, plus every active sequence's KV cache: bytes = W + B × KV_seq
- Compute ≈ 2 FLOPs per active parameter per sequence: FLOPs = 2 × P_active × B
Step time is set by whichever takes longer:
step time ≈ max( (W + B × KV_seq) ÷ (bandwidth × 0.7), 2 × P_active × B ÷ (FLOPS × 0.5) )
The 0.7 and 0.5 are typical efficiency factors for well-tuned engines. Then:
- Per-user speed = 1 ÷ step time
- Replica throughput = B ÷ step time
This is a first-order model. It ignores prefill load, attention kernel details and scheduling overhead, and lands within about 2× of real benchmarks. Use real benchmarks when you have them, and this estimate when you don't.
Worked example
Demand: a chat product peaking at 350 QPS, 2,000 input and 400 output tokens, so 140K output tokens/s at peak. The target is ≥ 30 tokens/s per user.
Replica shape: Llama 3.1 70B in FP8, which is 70.6 GB of weights. One H100 would leave under 10 GB for KV cache, so use TP = 2 (160 GB).
Batch size: KV per sequence = 2,400 tokens × 320 KB ≈ 0.79 GB. Try B = 64:
- Bytes per step: 70.6 GB + 64 × 0.79 GB ≈ 121 GB
- Bandwidth: 2 × 3.35 TB/s × 0.7 ≈ 4.7 TB/s, so memory time ≈ 25.8 ms
- Compute: 2 × 70.6B × 64 ÷ (2 × 990 TFLOPS × 0.5) ≈ 9 ms, so this is memory-bound
- Per-user speed: 1 ÷ 25.8 ms ≈ 39 tokens/s, which meets the target
- Memory check: 121 GB of 144 GB usable, so it fits
Replica throughput: 64 ÷ 25.8 ms ≈ 2,480 tokens/s
Fleet: keep 30% headroom for prefill work and bursts, which leaves ≈ 1,740 usable tokens/s per replica.
- Replicas = 140,000 ÷ 1,740 ≈ 81, so 162 H100s
- Cost ≈ 162 × 730 h × $2.50 ≈ $296K a month, always on
Back-of-envelope capacity estimator
From users to QPS to tokens per second, then to an API bill or a GPU count. The interview version of this is five lines on a whiteboard.
1. Traffic
2a. Buy: hosted API
2b. Build: self-hosted open model
Concurrent sequences decoding together.
Self-hosted numbers come from a roofline model (memory bandwidth vs compute) at 70% bandwidth and 50% compute efficiency. Real engines land within about 2× of this. That is good enough to pick between 10 and 100 GPUs, which is the decision an interview cares about.
Don't forget prefill
Prefill needs compute: 2 × P × input tokens. At peak that is 700K input tokens/s × 2 × 70.6B ≈ 99 PFLOPs per second, or ≈ 200 H100s at 50% efficiency if prefill ran on its own GPUs. That is a lot! In practice:
- Prefix caching removes much of it. Chat history and system prompts repeat across turns, and hit rates of 50–80% are common in multi-turn chat.
- Prefill is compute-bound, so it fills the compute that memory-bound decode leaves idle. That is what the headroom above is for.
- If prompts are very long (RAG with 20K tokens of context), size a separate prefill pool or at least double-check the numbers. Long-context products are often prefill-dominated.
Build vs buy
| Hosted API | Self-hosted | |
|---|---|---|
| Cost at low volume | Low: pay per token | High: GPUs idle but billed |
| Cost at high, steady volume | High | Often 3–10× cheaper per token |
| Model quality | Frontier models available | Open models only, a generation or so behind at the top end |
| Ops burden | None | GPUs, drivers, engines, autoscaling, on-call |
| Data control | Leaves your boundary (contracts, regions) | Fully yours |
| Elasticity | Instant, within rate limits | Minutes to scale, capacity-constrained |
Break-even check: self-hosting is cheaper when monthly API spend exceeds the fleet cost plus engineering time. The fleet must be sized for peak, so what matters is average utilisation. A fleet at 30% average utilisation pays more than three times the ideal cost per token. Blended strategies are common: self-host the high-volume, predictable workload and burst to an API for spikes and hard queries.
Key takeaways
- Start from peak output tokens per second, a TPOT target and a typical context length.
- Per replica: choose GPUs so the weights fit with room for KV cache, pick the largest batch that meets TPOT, then throughput = batch ÷ step time.
- Replicas = peak tokens/s ÷ (replica throughput × utilisation headroom). Then check memory and prefill load separately.
- Self-hosting only wins with high, steady utilisation. Idle GPUs cost the same as busy ones.