GenAI System Design
2. Inference, GPUs and Serving

Capacity Planning for Inference

A step-by-step method to size a self-hosted LLM fleet, from traffic to GPUs, with the roofline estimate behind it and a build vs buy comparison.

Lesson 7 of 7 13 min

The method

1. Demand
peak output tokens/s, context, TPOT target
2. Replica shape
model, precision, GPUs per replica
3. Batch size
largest that meets TPOT and fits KV
4. Replica throughput
batch ÷ step time
5. Fleet
replicas, GPUs, $/month

The roofline estimate behind it

For each decode step, the GPU must:

  • Read all weights once, plus every active sequence's KV cache: bytes = W + B × KV_seq
  • Compute ≈ 2 FLOPs per active parameter per sequence: FLOPs = 2 × P_active × B

Step time is set by whichever takes longer:

step time ≈ max( (W + B × KV_seq) ÷ (bandwidth × 0.7), 2 × P_active × B ÷ (FLOPS × 0.5) )

The 0.7 and 0.5 are typical efficiency factors for well-tuned engines. Then:

  • Per-user speed = 1 ÷ step time
  • Replica throughput = B ÷ step time

This is a first-order model. It ignores prefill load, attention kernel details and scheduling overhead, and lands within about 2× of real benchmarks. Use real benchmarks when you have them, and this estimate when you don't.

Worked example

Demand: a chat product peaking at 350 QPS, 2,000 input and 400 output tokens, so 140K output tokens/s at peak. The target is ≥ 30 tokens/s per user.

Replica shape: Llama 3.1 70B in FP8, which is 70.6 GB of weights. One H100 would leave under 10 GB for KV cache, so use TP = 2 (160 GB).

Batch size: KV per sequence = 2,400 tokens × 320 KB ≈ 0.79 GB. Try B = 64:

  • Bytes per step: 70.6 GB + 64 × 0.79 GB ≈ 121 GB
  • Bandwidth: 2 × 3.35 TB/s × 0.7 ≈ 4.7 TB/s, so memory time ≈ 25.8 ms
  • Compute: 2 × 70.6B × 64 ÷ (2 × 990 TFLOPS × 0.5) ≈ 9 ms, so this is memory-bound
  • Per-user speed: 1 ÷ 25.8 ms ≈ 39 tokens/s, which meets the target
  • Memory check: 121 GB of 144 GB usable, so it fits

Replica throughput: 64 ÷ 25.8 ms ≈ 2,480 tokens/s

Fleet: keep 30% headroom for prefill work and bursts, which leaves ≈ 1,740 usable tokens/s per replica.

  • Replicas = 140,000 ÷ 1,740 ≈ 81, so 162 H100s
  • Cost ≈ 162 × 730 h × $2.50 ≈ $296K a month, always on

Back-of-envelope capacity estimator

From users to QPS to tokens per second, then to an API bill or a GPU count. The interview version of this is five lines on a whiteboard.

×

1. Traffic

Requests per day
10M
Average QPS
115.7
requests/day ÷ 86,400
Peak QPS
347.2
average × 3
Peak output tokens/s
138.9K
peak QPS × output tokens

2a. Buy: hosted API

Tokens per month
720B
600B in, 120B out
API cost per month
$2,400,000
list price, no caching or batch discount

2b. Build: self-hosted open model

Concurrent sequences decoding together.

GPUs per replica
2
batch 64
Per-user speed
39 tok/s
memory-bound decode
Replicas at peak
80 × 2 = 160 GPUs
2.5K tok/s each, 30% headroom
GPU cost per month
$292,000
at ~$2.5/GPU-hour, always on

Self-hosted numbers come from a roofline model (memory bandwidth vs compute) at 70% bandwidth and 50% compute efficiency. Real engines land within about 2× of this. That is good enough to pick between 10 and 100 GPUs, which is the decision an interview cares about.

Don't forget prefill

Prefill needs compute: 2 × P × input tokens. At peak that is 700K input tokens/s × 2 × 70.6B ≈ 99 PFLOPs per second, or ≈ 200 H100s at 50% efficiency if prefill ran on its own GPUs. That is a lot! In practice:

  • Prefix caching removes much of it. Chat history and system prompts repeat across turns, and hit rates of 50–80% are common in multi-turn chat.
  • Prefill is compute-bound, so it fills the compute that memory-bound decode leaves idle. That is what the headroom above is for.
  • If prompts are very long (RAG with 20K tokens of context), size a separate prefill pool or at least double-check the numbers. Long-context products are often prefill-dominated.

Build vs buy

Hosted APISelf-hosted
Cost at low volumeLow: pay per tokenHigh: GPUs idle but billed
Cost at high, steady volumeHighOften 3–10× cheaper per token
Model qualityFrontier models availableOpen models only, a generation or so behind at the top end
Ops burdenNoneGPUs, drivers, engines, autoscaling, on-call
Data controlLeaves your boundary (contracts, regions)Fully yours
ElasticityInstant, within rate limitsMinutes to scale, capacity-constrained

Break-even check: self-hosting is cheaper when monthly API spend exceeds the fleet cost plus engineering time. The fleet must be sized for peak, so what matters is average utilisation. A fleet at 30% average utilisation pays more than three times the ideal cost per token. Blended strategies are common: self-host the high-volume, predictable workload and burst to an API for spikes and hard queries.

Key takeaways

  • Start from peak output tokens per second, a TPOT target and a typical context length.
  • Per replica: choose GPUs so the weights fit with room for KV cache, pick the largest batch that meets TPOT, then throughput = batch ÷ step time.
  • Replicas = peak tokens/s ÷ (replica throughput × utilisation headroom). Then check memory and prefill load separately.
  • Self-hosting only wins with high, steady utilisation. Idle GPUs cost the same as busy ones.

Go deeper