Back-of-Envelope Math for LLM Systems
Turn a user count into QPS, tokens per second, concurrent sequences, GPUs and a monthly bill, with one worked example you can reuse in any interview.
The estimation chain
Every LLM capacity question reduces to the same chain. Memorise it once:
Tokens per second is the unit that matters. QPS alone is misleading: 100 QPS of short classifications and 100 QPS of 2,000-token essays differ in cost by about 100×.
Worked example: a consumer chat assistant
Assumptions, stated out loud and rounded:
- 1M daily active users, 10 messages each per day
- 2,000 input tokens per request (system prompt, history, retrieved context) and 400 output tokens
- Peak is 3× the daily average
Traffic
- Requests per day: 1M × 10 = 10M
- Average QPS: 10M ÷ 100K s ≈ 100 QPS (86,400 s gives 116; either is fine)
- Peak QPS: 3 × ≈ 116 ≈ 350 QPS
Tokens
- Peak output: 350 × 400 = 140K tokens/s
- Peak input (prefill): 350 × 2,000 = 700K tokens/s
- Per month: 10M × 30 × 2,400 ≈ 720B tokens
Option A: buy (hosted API at $2 in / $10 out per million)
- Input: 600B × $2/1M = $1.2M
- Output: 120B × $10/1M = $1.2M
- ≈ $2.4M a month at list price, before caching discounts
Option B: build (self-host a 70B open model in FP8 on H100s)
- A 70B model in FP8 is ≈ 70 GB of weights, so we use 2 × H100 per replica for weights plus KV cache.
- At a batch of about 64 sequences, one replica decodes roughly 2,500 tokens/s in total, about 40 tokens/s per user. The capacity planning lesson derives this number.
- Keep 30% headroom for prefill and bursts, which leaves ≈ 1,750 usable tokens/s per replica.
- Replicas: 140K ÷ 1,750 = 80 replicas = 160 H100s
- At ≈ $2.50 per GPU-hour: 160 × 730 h × $2.50 ≈ $290K a month
So at this scale self-hosting looks about 8× cheaper, if a 70B open model meets your quality bar and you can run a GPU fleet reliably. At 1/100th of this traffic, you would need only one or two replicas. They would sit mostly idle, and the API wins easily.
Concurrency with Little's law
GPUs serve many users at once, so you also need to know how many sequences are in flight:
Concurrent requests = arrival rate × time each request takes
Each request takes ≈ 0.5 s TTFT + 400 tokens ÷ 40 tokens/s = ≈ 10.5 s. At 350 QPS: 350 × 10.5 ≈ 3,700 concurrent sequences. At 64 per replica, that is ≈ 58 replicas, the same order of magnitude as the throughput estimate. When two independent estimates agree, say so. It shows the number is not a fluke.
Try it
Back-of-envelope capacity estimator
From users to QPS to tokens per second, then to an API bill or a GPU count. The interview version of this is five lines on a whiteboard.
1. Traffic
2a. Buy: hosted API
2b. Build: self-hosted open model
Concurrent sequences decoding together.
Self-hosted numbers come from a roofline model (memory bandwidth vs compute) at 70% bandwidth and 50% compute efficiency. Real engines land within about 2× of this. That is good enough to pick between 10 and 100 GPUs, which is the decision an interview cares about.
Storage estimates still matter
- Conversation history: 10M requests × ≈ 2.4K tokens × ≈ 4 bytes per token ≈ 100 GB of text per day, or about 36 TB a year before compression.
- Embeddings for RAG: 10M document chunks × 1,536 dims × 4 bytes ≈ 61 GB of raw vectors. See vector index sizing.
Key takeaways
- The chain: DAU → requests/day → peak QPS → tokens/s → either API dollars or GPU count.
- A day is ≈ 100K seconds. Peak traffic is typically 2–5× the daily average.
- Little's law gives concurrency: QPS × seconds per request = sequences the GPUs must hold at once.
- Size self-hosted fleets from output tokens per second per replica, and keep headroom for prefill and bursts.