Numbers Every AI Engineer Should Know
The latency, token, memory and hardware numbers that let you sanity-check any GenAI design, and three derivations that show how they connect.
Why memorise numbers
Classic system design has Jeff Dean's "latency numbers every programmer should know". GenAI needs its own set. With them you can spot an impossible requirement quickly. Asking for "sub-100 ms responses from a 70B model" is one example. You can also justify a design without a calculator.
Tokens
| 1 token (English) | ≈ 4 characters ≈ 0.75 words | Converts product requirements (pages, messages) into tokens. |
|---|---|---|
| 1 page of text | ≈ 500–700 tokens | Sizing documents for RAG ingestion and context budgets. |
| Human reading speed | ≈ 5–8 tokens/s | Streaming faster than ~20 tokens/s feels instant to readers. |
| Chat turn | ≈ 1–3K tokens in, 200–500 out | Default for traffic estimates when nobody gives you numbers. |
Latency
| Time to first token (hosted API) | ≈ 0.2–1 s | Grows with prompt length and model size. |
|---|---|---|
| Decode speed (hosted frontier model) | ≈ 50–150 tokens/s | A 500-token answer takes 3–10 s end to end. |
| Embedding a query (API) | ≈ 50–300 ms | Self-hosted small embedders: 5–20 ms. |
| HNSW search, 10M vectors in RAM | ≈ 1–10 ms | Vector search is rarely the latency bottleneck; the LLM is. |
| Cross-encoder rerank of 50 docs (GPU) | ≈ 50–150 ms | Worth it for quality, but budget for it. |
Hardware
| H100 SXM | 80 GB HBM3, 3.35 TB/s, ~990 TFLOPS BF16 | The reference GPU for capacity estimates. |
|---|---|---|
| H200 / B200 | 141 GB, 4.8 TB/s / 180 GB, 8 TB/s | More memory and bandwidth mean bigger batches and faster decode. |
| NVLink (H100) | 900 GB/s per GPU | Why tensor parallelism stays inside one server. |
| InfiniBand NDR | 400 Gb/s ≈ 50 GB/s per link | Cross-node links are ~18× slower than NVLink. |
| H100 rental | ≈ $2–4 per GPU-hour | ≈ $1.5–3K per GPU-month. Commitments cut this substantially. |
Memory
| Model weights (BF16) | 2 bytes × parameters | 70B model ≈ 140 GB, so it needs at least 2 × 80 GB GPUs. |
|---|---|---|
| Model weights (INT4) | ≈ 0.5–0.6 bytes × parameters | 70B fits on one 80 GB GPU with room for KV cache. |
| KV cache, Llama 3.1 70B | ≈ 320 KB per token (BF16) | A 32K-token conversation holds ≈ 10 GB of GPU memory. |
| KV cache, Llama 3.1 8B | ≈ 128 KB per token (BF16) | Formula: 2 × layers × kv_heads × head_dim × bytes. |
| 1536-dim float32 embedding | 6 KB | 100M vectors ≈ 600 GB raw, before index overhead. |
Figures checked 2026-09-28. Hardware prices move fast; the ratios between them move slowly.
Three derivations worth knowing cold
1. How fast can a GPU generate tokens?
Each decode step reads every weight once to produce one token per sequence. At small batch sizes the step time is set by memory bandwidth:
tokens/s per sequence ≈ memory bandwidth ÷ weight bytes
- Llama 3.1 8B in BF16 = 16 GB. On an H100 (3.35 TB/s): 3,350 ÷ 16 ≈ 210 tokens/s ceiling. Real engines reach 60–80% of this.
- Llama 3.1 70B in FP8 = 70 GB on one H100: 3,350 ÷ 70 ≈ 48 tokens/s.
That is why quantization and faster memory (H200, B200) speed up generation. They reduce the bytes each step must read, or read them faster.
2. How many GPUs just to hold the model?
Weight memory = parameters × bytes per parameter
| Model | BF16 (2 B) | FP8 (1 B) | INT4 (≈ 0.55 B) |
|---|---|---|---|
| 8B | 16 GB | 8 GB | ≈ 4.5 GB |
| 70B | 140 GB | 70 GB | ≈ 39 GB |
| 405B | 810 GB | 405 GB | ≈ 225 GB |
Then add room for the KV cache and runtime. A useful rule is to keep weights under about 60–70% of total GPU memory for a serving workload.
3. How much memory does one conversation hold?
KV cache per token = 2 (K and V) × layers × KV heads × head dim × bytes per value
For Llama 3.1 70B: 2 × 80 × 8 × 128 × 2 bytes = 327,680 bytes ≈ 320 KB per token.
- One 32K-token conversation takes ≈ 10.7 GB.
- 100 users at 8K tokens each take ≈ 268 GB, more than the model weights.
Where time goes in a RAG request
| Step | Typical time |
|---|---|
| Embed the query | 10–100 ms |
| Vector search (HNSW, in memory) | 1–10 ms |
| Keyword search (BM25) | 5–30 ms |
| Rerank top 50 with a cross-encoder | 50–150 ms |
| LLM time to first token | 200–1,000 ms |
| LLM generation of 400 tokens | 3–8 s |
The LLM dominates. Candidates sometimes spend ten minutes optimising the vector database for a system where it accounts for 0.1% of latency. Show the table, then spend your time where the seconds are.
Key takeaways
- Decode speed ceiling ≈ memory bandwidth ÷ bytes of weights. For an 8B BF16 model on an H100 that is ≈ 200 tokens/s.
- Weights take 2 bytes per parameter in BF16. A 70B model needs ≈ 140 GB, which means two 80 GB GPUs.
- KV cache per token = 2 × layers × KV heads × head dim × bytes. Llama 3.1 70B uses ≈ 320 KB per token.
- Vector search takes milliseconds while the LLM takes seconds. Optimise the LLM first.