GenAI System Design
0. How to Ace a GenAI System Design Interview

Numbers Every AI Engineer Should Know

The latency, token, memory and hardware numbers that let you sanity-check any GenAI design, and three derivations that show how they connect.

Lesson 5 of 6 8 min

Why memorise numbers

Classic system design has Jeff Dean's "latency numbers every programmer should know". GenAI needs its own set. With them you can spot an impossible requirement quickly. Asking for "sub-100 ms responses from a 70B model" is one example. You can also justify a design without a calculator.

Tokens

1 token (English)≈ 4 characters ≈ 0.75 wordsConverts product requirements (pages, messages) into tokens.
1 page of text≈ 500–700 tokensSizing documents for RAG ingestion and context budgets.
Human reading speed≈ 5–8 tokens/sStreaming faster than ~20 tokens/s feels instant to readers.
Chat turn≈ 1–3K tokens in, 200–500 outDefault for traffic estimates when nobody gives you numbers.

Latency

Time to first token (hosted API)≈ 0.2–1 sGrows with prompt length and model size.
Decode speed (hosted frontier model)≈ 50–150 tokens/sA 500-token answer takes 3–10 s end to end.
Embedding a query (API)≈ 50–300 msSelf-hosted small embedders: 5–20 ms.
HNSW search, 10M vectors in RAM≈ 1–10 msVector search is rarely the latency bottleneck; the LLM is.
Cross-encoder rerank of 50 docs (GPU)≈ 50–150 msWorth it for quality, but budget for it.

Hardware

H100 SXM80 GB HBM3, 3.35 TB/s, ~990 TFLOPS BF16The reference GPU for capacity estimates.
H200 / B200141 GB, 4.8 TB/s / 180 GB, 8 TB/sMore memory and bandwidth mean bigger batches and faster decode.
NVLink (H100)900 GB/s per GPUWhy tensor parallelism stays inside one server.
InfiniBand NDR400 Gb/s ≈ 50 GB/s per linkCross-node links are ~18× slower than NVLink.
H100 rental≈ $2–4 per GPU-hour≈ $1.5–3K per GPU-month. Commitments cut this substantially.

Memory

Model weights (BF16)2 bytes × parameters70B model ≈ 140 GB, so it needs at least 2 × 80 GB GPUs.
Model weights (INT4)≈ 0.5–0.6 bytes × parameters70B fits on one 80 GB GPU with room for KV cache.
KV cache, Llama 3.1 70B≈ 320 KB per token (BF16)A 32K-token conversation holds ≈ 10 GB of GPU memory.
KV cache, Llama 3.1 8B≈ 128 KB per token (BF16)Formula: 2 × layers × kv_heads × head_dim × bytes.
1536-dim float32 embedding6 KB100M vectors ≈ 600 GB raw, before index overhead.

Figures checked 2026-09-28. Hardware prices move fast; the ratios between them move slowly.

Three derivations worth knowing cold

1. How fast can a GPU generate tokens?

Each decode step reads every weight once to produce one token per sequence. At small batch sizes the step time is set by memory bandwidth:

tokens/s per sequence ≈ memory bandwidth ÷ weight bytes

  • Llama 3.1 8B in BF16 = 16 GB. On an H100 (3.35 TB/s): 3,350 ÷ 16 ≈ 210 tokens/s ceiling. Real engines reach 60–80% of this.
  • Llama 3.1 70B in FP8 = 70 GB on one H100: 3,350 ÷ 70 ≈ 48 tokens/s.

That is why quantization and faster memory (H200, B200) speed up generation. They reduce the bytes each step must read, or read them faster.

2. How many GPUs just to hold the model?

Weight memory = parameters × bytes per parameter

ModelBF16 (2 B)FP8 (1 B)INT4 (≈ 0.55 B)
8B16 GB8 GB≈ 4.5 GB
70B140 GB70 GB≈ 39 GB
405B810 GB405 GB≈ 225 GB

Then add room for the KV cache and runtime. A useful rule is to keep weights under about 60–70% of total GPU memory for a serving workload.

3. How much memory does one conversation hold?

KV cache per token = 2 (K and V) × layers × KV heads × head dim × bytes per value

For Llama 3.1 70B: 2 × 80 × 8 × 128 × 2 bytes = 327,680 bytes ≈ 320 KB per token.

  • One 32K-token conversation takes ≈ 10.7 GB.
  • 100 users at 8K tokens each take ≈ 268 GB, more than the model weights.

Where time goes in a RAG request

StepTypical time
Embed the query10–100 ms
Vector search (HNSW, in memory)1–10 ms
Keyword search (BM25)5–30 ms
Rerank top 50 with a cross-encoder50–150 ms
LLM time to first token200–1,000 ms
LLM generation of 400 tokens3–8 s

The LLM dominates. Candidates sometimes spend ten minutes optimising the vector database for a system where it accounts for 0.1% of latency. Show the table, then spend your time where the seconds are.

Key takeaways

  • Decode speed ceiling ≈ memory bandwidth ÷ bytes of weights. For an 8B BF16 model on an H100 that is ≈ 200 tokens/s.
  • Weights take 2 bytes per parameter in BF16. A 70B model needs ≈ 140 GB, which means two 80 GB GPUs.
  • KV cache per token = 2 × layers × KV heads × head dim × bytes. Llama 3.1 70B uses ≈ 320 KB per token.
  • Vector search takes milliseconds while the LLM takes seconds. Optimise the LLM first.

Go deeper