A Worked Estimate, End to End
A complete capacity and cost estimate for an AI-native product, covering traffic, tokens, retrieval, vector storage, GPUs or API spend, and cost per user, as you would present it in an interview.
The problem
Design capacity for an AI assistant inside a document-collaboration product: 5M monthly active users, 1M daily active. Users ask questions about their workspace documents, and answers need citations.
1. Assumptions (say them out loud)
| Assumption | Value |
|---|---|
| DAU | 1M |
| Questions per DAU per day | 8 |
| Peak ÷ average | 3× |
| Input tokens per request | 1.5K system and history + 3K retrieved = 4.5K |
| Output tokens | 350 |
| Prefix cached (system prompt + tools) | 1.2K of the 4.5K |
| Documents per MAU | 200 docs × 3 pages ≈ 600 pages |
2. Traffic
- Requests per day: 1M × 8 = 8M
- Average QPS: 8M ÷ 86,400 ≈ 93, peak ≈ 280 QPS
- Peak output tokens/s: 280 × 350 ≈ 98K tokens/s
- Peak input tokens/s: 280 × 4.5K ≈ 1.26M tokens/s
3. Model cost (buy)
On a mid-tier API at $2 / $10 per million, $0.20 cached:
- Fresh input: 8M × 3.3K = 26.4B tokens a day × $2/1M = $52.8K a day
- Cached input: 8M × 1.2K = 9.6B × $0.20/1M = $1.9K a day
- Output: 8M × 350 = 2.8B × $10/1M = $28K a day
- ≈ $83K a day, ≈ $2.5M a month
Levers: route simple questions (say 50%) to a small model, trim retrieved context from 3K to 2K tokens with better reranking, and shorter answers. Together these can plausibly cut the bill by 2–3×.
4. Model cost (build), cross-checked
Self-host a 70B-class model in FP8 on 2 × H100 replicas. From Module 2's method, ≈ 2,000–2,500 output tokens/s per replica at interactive TPOT, but prefill is heavy here (4.5K-token prompts), so assume ≈ 1,500 usable output tokens/s per replica after headroom.
- Replicas ≈ 98K ÷ 1,500 ≈ 65, so 130 H100s
- Cost: 130 × $2.50 × 730 ≈ $240K a month, plus engineering
Cross-check with Little's law: each request lasts ≈ 0.8 s TTFT + 350 ÷ 35 tokens/s ≈ 11 s. So 280 QPS × 11 s ≈ 3,100 concurrent sequences. KV per sequence ≈ 4.85K tokens × 320 KB ≈ 1.55 GB. At about 45 sequences per replica (≈ 70 GB free KV memory on 2 × H100 in FP8), that is ≈ 69 replicas, consistent with 65.
5. Retrieval infrastructure
- Pages: 5M MAU × 600 = 3B pages, which is too many to embed naively. Embed on first access or for active workspaces. Suppose 1B chunks are indexed.
- Vectors: 1B × 1,024 dims. At float32 that is 4 TB. With int8 it is 1 TB, or with PQ at 64 bytes per vector, 64 GB plus full vectors on SSD for rescoring.
- The index is partitioned by workspace (the tenant), so each query searches only its own workspace's vectors. Most workspaces are small, which makes search cheap.
- Embedding cost: 1B chunks × 400 tokens = 400B tokens. At $0.02 per million that is ≈ $8K one-off, plus ongoing edits.
- Search QPS: 280 peak, which a modest cluster handles easily. Retrieval cost is small compared with the LLM, likely under 5% of the total.
6. Storage
- Conversation history: 8M requests × ≈ 5K tokens × 4 bytes ≈ 160 GB a day raw, so keep 30 days hot (≈ 5 TB), compress, and archive or delete per policy.
- Traces: sample content at 5%, and keep metrics at 100%.
7. Cost per user
- API path: $2.5M ÷ 5M MAU ≈ $0.50 per MAU per month (≈ $2.50 per DAU).
- With the optimisations: ≈ $0.20–0.25 per MAU.
Compare with pricing. If the assistant is part of a $10-per-seat plan, it works. If it is free for all users, it needs limits or a paid tier.
8. Sensitivities
| If… | Effect |
|---|---|
| Retrieved context grows 3K → 6K tokens | Input cost roughly doubles |
| Questions per DAU go 8 → 20 | Everything ×2.5 |
| Peak factor 3 → 5 | Self-hosted fleet ×1.7, API bill unchanged |
| A small model handles 50% of traffic | Model bill ≈ −40% |
Key takeaways
- Estimate in layers, from traffic to tokens to the model bill (API or GPUs) to retrieval infrastructure to storage, then cost per user.
- State assumptions, round aggressively, and cross-check with a second method, such as Little's law against throughput.
- The LLM usually dominates cost. Retrieval and storage are real but typically much smaller.
- Finish with sensitivities and the levers that change the answer most.