GenAI System Design
10. GenAI System Design Problems

Designing a ChatGPT-Style Chat Service

A full walkthrough for a consumer AI chat product with 10M daily users, covering streaming, conversation storage, context management, GPU capacity, prefix caching, plan-based limits and overload handling.

Lesson 1 of 10 16 min

1. Requirements

Functional

  • Multi-turn chat with streamed responses, and several models (fast and advanced)
  • Conversation history: list, resume, rename, delete
  • File and image uploads (in scope for context, but covered briefly)
  • Plans: free with limits, paid with higher limits and better models

Non-functional

  • 10M DAU, 15 messages per user per day
  • First token at p95 under 1 s. Streaming at 30 tokens/s or more
  • High availability. Degrade, don't fail, under load spikes
  • Safety moderation, and privacy (users can delete conversations)

2. Operating point

Consumer chat is latency-sensitive (TTFT) and cost-dominated. Quality matters, but most messages are everyday tasks. So: stream everything, default to an efficient model, give paid users the advanced model, and keep context tight.

3. Estimates

QuantityValue
Messages per day10M × 15 = 150M
Average / peak QPS (3×)≈ 1,740 / ≈ 5,200
Input tokens per turn≈ 3K (system prompt + history + message)
Output tokens per turn≈ 500
Peak output tokens/s5,200 × 500 ≈ 2.6M tokens/s
Peak input tokens/s5,200 × 3K ≈ 15.6M tokens/s (mostly cached prefix)
Concurrent streams (Little's law)5,200 × ≈ 15 s ≈ 78K
Conversation text stored per day150M × ≈ 600 tokens (message + reply) × 4 B ≈ 360 GB (≈ 130 TB per year raw)

GPU capacity, assuming we serve a 70B-class model ourselves: ≈ 1,500 usable output tokens/s per 2-GPU replica gives 2.6M ÷ 1,500 ≈ 1,700 replicas, or ≈ 3,500 GPUs at peak for this one model. Every percentage point of efficiency is worth tens of GPUs, which is why the deep dives focus on caching and batching.

4. Architecture

Client
web / mobile, SSE
Edge + API gateway
auth, plan limits
Chat service
conversation state
Context builder
history, memory, files
Model router
plan, model, affinity
Inference pools
vLLM/SGLang per model

Supporting systems

  • Conversation store: messages sharded by user_id (Cassandra, DynamoDB or sharded Postgres), with recent conversations cached.
  • Object storage for uploads, plus a parsing pipeline (text extraction, chunking, and embeddings for long files).
  • Moderation service running in parallel on input and on streamed output.
  • Async workers: conversation titles, memory extraction, summaries, analytics, all off the critical path.
  • Usage metering: token counts per user from the stream's final usage event, feeding limits and billing.

Request flow: authenticate, check plan limits, then load the conversation. The context builder assembles the prompt: system prompt, memory, summary of older turns, recent turns, and the new message. The router picks the model and a replica (with session affinity), and the response streams back over SSE while the chat service appends it to the store when the stream completes.

5. Deep dives

Deep dive A: inference efficiency

  • Prefix caching plus affinity: turn n's prompt is turn n−1's prompt plus the last answer plus the new message. Route a conversation to the same replica (a consistent hash on conversation ID) so its KV prefix is warm. Cache hit rates of 70–90% on input tokens are realistic, which cuts prefill compute and TTFT dramatically.
  • Continuous batching at a batch size tuned to the TPOT target (≈ 30 tokens/s per user).
  • Model tiering: most free-tier traffic goes to a smaller model. That changes the GPU count more than any other lever.
  • FP8 weights and KV cache to double effective KV capacity and fit more concurrent streams per replica.
  • Separate pools for long-context conversations, so they don't shrink batch sizes for everyone else.

Deep dive B: context management

Conversations can run to hundreds of turns. Policy:

  1. Always include the system prompt and the user's saved memory (a few hundred tokens).
  2. Include recent turns verbatim, up to a budget (for example 6K tokens).
  3. Replace older turns with a rolling summary, regenerated asynchronously every N turns.
  4. For uploaded documents, retrieve relevant chunks rather than including whole files.

This keeps average input near 3K tokens instead of growing without bound. It trades some recall of early details for bounded cost and latency.

Deep dive C: overload and fairness

At a peak or a viral moment:

  • Per-plan token limits (sliding window), enforced at the gateway before GPUs are touched.
  • Admission control: when queue time would push TTFT past the SLO, shed or queue free-tier requests first, and route them to a smaller model ("high demand, using a faster model").
  • Autoscaling on queue depth and KV usage, with warm spare capacity, because loading weights takes minutes.
  • Priority lanes: paid and interactive traffic ahead of background tasks such as titles and summaries.

6. Evaluation and safety

  • Offline: an eval suite per model version (helpfulness, instruction following, safety refusals, regressions) before any model swap.
  • Online: thumbs, regenerate rate, conversation length, retention, and A/B tests for model or prompt changes.
  • Safety: input and output moderation classifiers, and abuse detection (scripted accounts, jailbreak farms).
  • Privacy: deleting a conversation removes it from the store, caches and memory, and it is excluded from training datasets per the user's settings.

7. Bottlenecks and follow-ups

Likely questionAnswer sketch
"GPUs run out at peak?"Tier models, queue free traffic, burst to another region or provider, use off-peak for batch work
"A user sends a 200K-token file?"Retrieval over chunks, per-plan context caps, a dedicated long-context pool
"Make it cheaper by 2×?"Higher cache hit rate, smaller default model, FP8, shorter default answers
"Multi-region?"Stateless services everywhere, conversation store replicated by region of residence, GPU pools per region with failover

Key takeaways

  • At consumer scale the system is GPU-bound. Capacity is planned in output tokens per second at peak, and prefix caching of conversation history is essential.
  • Context management (truncation, summaries, memory) keeps per-turn tokens bounded as conversations grow.
  • Route conversations to replicas with affinity, so each turn reuses the cached prefix of the last.
  • Degrade gracefully under load by queueing, switching free users to smaller models, and enforcing per-plan token limits.

Go deeper

Finished reading? Mark it done to track your progress.