GenAI System Design
0. How to Ace a GenAI System Design Interview

Back-of-Envelope Math for LLM Systems

Turn a user count into QPS, tokens per second, concurrent sequences, GPUs and a monthly bill, with one worked example you can reuse in any interview.

Lesson 4 of 6 12 min

The estimation chain

Every LLM capacity question reduces to the same chain. Memorise it once:

Users
DAU × requests/user
Requests/day
÷ 100K s
Peak QPS
× peak factor
Tokens/s
× tokens per request
API $ or GPUs
price, or tokens/s per GPU

Tokens per second is the unit that matters. QPS alone is misleading: 100 QPS of short classifications and 100 QPS of 2,000-token essays differ in cost by about 100×.

Worked example: a consumer chat assistant

Assumptions, stated out loud and rounded:

  • 1M daily active users, 10 messages each per day
  • 2,000 input tokens per request (system prompt, history, retrieved context) and 400 output tokens
  • Peak is 3× the daily average

Traffic

  • Requests per day: 1M × 10 = 10M
  • Average QPS: 10M ÷ 100K s ≈ 100 QPS (86,400 s gives 116; either is fine)
  • Peak QPS: 3 × ≈ 116 ≈ 350 QPS

Tokens

  • Peak output: 350 × 400 = 140K tokens/s
  • Peak input (prefill): 350 × 2,000 = 700K tokens/s
  • Per month: 10M × 30 × 2,400 ≈ 720B tokens

Option A: buy (hosted API at $2 in / $10 out per million)

  • Input: 600B × $2/1M = $1.2M
  • Output: 120B × $10/1M = $1.2M
  • ≈ $2.4M a month at list price, before caching discounts

Option B: build (self-host a 70B open model in FP8 on H100s)

  • A 70B model in FP8 is ≈ 70 GB of weights, so we use 2 × H100 per replica for weights plus KV cache.
  • At a batch of about 64 sequences, one replica decodes roughly 2,500 tokens/s in total, about 40 tokens/s per user. The capacity planning lesson derives this number.
  • Keep 30% headroom for prefill and bursts, which leaves ≈ 1,750 usable tokens/s per replica.
  • Replicas: 140K ÷ 1,750 = 80 replicas = 160 H100s
  • At ≈ $2.50 per GPU-hour: 160 × 730 h × $2.50 ≈ $290K a month

So at this scale self-hosting looks about 8× cheaper, if a 70B open model meets your quality bar and you can run a GPU fleet reliably. At 1/100th of this traffic, you would need only one or two replicas. They would sit mostly idle, and the API wins easily.

Concurrency with Little's law

GPUs serve many users at once, so you also need to know how many sequences are in flight:

Concurrent requests = arrival rate × time each request takes

Each request takes ≈ 0.5 s TTFT + 400 tokens ÷ 40 tokens/s = ≈ 10.5 s. At 350 QPS: 350 × 10.5 ≈ 3,700 concurrent sequences. At 64 per replica, that is ≈ 58 replicas, the same order of magnitude as the throughput estimate. When two independent estimates agree, say so. It shows the number is not a fluke.

Try it

Back-of-envelope capacity estimator

From users to QPS to tokens per second, then to an API bill or a GPU count. The interview version of this is five lines on a whiteboard.

×

1. Traffic

Requests per day
10M
Average QPS
115.7
requests/day ÷ 86,400
Peak QPS
347.2
average × 3
Peak output tokens/s
138.9K
peak QPS × output tokens

2a. Buy: hosted API

Tokens per month
720B
600B in, 120B out
API cost per month
$2,400,000
list price, no caching or batch discount

2b. Build: self-hosted open model

Concurrent sequences decoding together.

GPUs per replica
2
batch 64
Per-user speed
39 tok/s
memory-bound decode
Replicas at peak
80 × 2 = 160 GPUs
2.5K tok/s each, 30% headroom
GPU cost per month
$292,000
at ~$2.5/GPU-hour, always on

Self-hosted numbers come from a roofline model (memory bandwidth vs compute) at 70% bandwidth and 50% compute efficiency. Real engines land within about 2× of this. That is good enough to pick between 10 and 100 GPUs, which is the decision an interview cares about.

Storage estimates still matter

  • Conversation history: 10M requests × ≈ 2.4K tokens × ≈ 4 bytes per token ≈ 100 GB of text per day, or about 36 TB a year before compression.
  • Embeddings for RAG: 10M document chunks × 1,536 dims × 4 bytes ≈ 61 GB of raw vectors. See vector index sizing.

Key takeaways

  • The chain: DAU → requests/day → peak QPS → tokens/s → either API dollars or GPU count.
  • A day is ≈ 100K seconds. Peak traffic is typically 2–5× the daily average.
  • Little's law gives concurrency: QPS × seconds per request = sequences the GPUs must hold at once.
  • Size self-hosted fleets from output tokens per second per replica, and keep headroom for prefill and bursts.

Go deeper