GenAI System Design
0. How to Ace a GenAI System Design Interview

How GenAI Design Rounds Differ

What stays the same and what changes when an LLM sits in the request path, and the three kinds of GenAI system design questions you will be asked.

Lesson 1 of 6 7 min

The same interview, with a new bottleneck

A GenAI system design round still rewards the habits you already have. You clarify requirements, estimate scale, draw a request path, and deep-dive into the part most likely to break. Load balancers, queues, caches and databases all still appear.

Three things change, and together they reshape most designs:

  1. The core operation is expensive. A cache hit costs microseconds and a fraction of a cent per million. A single LLM call takes 1 to 10 seconds, and on a frontier model it costs around a cent.
  2. The bottleneck is a GPU, and specifically its memory. Model weights and the per-user KV cache must fit in GPU memory. Generation speed is set by how fast that memory can be read, not by CPU cores.
  3. The output is probabilistic. The same input can produce different answers, some of them wrong. You cannot unit-test your way to correctness, so evaluation becomes part of the architecture.

Classic vs GenAI systems at a glance

DimensionClassic web serviceLLM-powered service
Typical latency10–200 ms1–10 s (streamed), first token in 0.2–1 s
Cost per request≈ $0.000001$0.001–$0.05
Scaling unitStateless pod on a CPU nodeGPU replica at $2–30 per hour
BottleneckCPU, DB connections, I/OGPU memory capacity and bandwidth
Per-request stateIn a databaseIn the context window and KV cache
CorrectnessDeterministic, testedProbabilistic, evaluated
New failure modesTimeouts, errorsHallucination, prompt injection, runaway agent loops

The AI-native request path

Most GenAI products share a skeleton. Learn it once, and each design question becomes a matter of which boxes matter most.

Client
streams tokens (SSE)
LLM gateway
auth, rate limits, routing
Orchestrator
prompt, tools, agent loop
Retrieval
embed, vector + keyword search
Model inference
GPU serving or hosted API
Guardrails
filter input and output
Online path. Offline pipelines (document ingestion, embedding, evals, fine-tuning) feed the retrieval and model boxes.

Behind the online path sit offline pipelines. They ingest and chunk documents, compute embeddings, run evaluation suites, and sometimes fine-tune models. In classic terms these are batch jobs, and they are often where freshness and cost problems hide.

Three kinds of questions

1. Build an LLM-powered product. Examples: a support chatbot over a company's docs, an AI coding assistant, meeting summaries. The core of the design is retrieval, prompt construction, evaluation and cost per user. Usually you call a hosted model.

2. Build LLM infrastructure. Examples: design a model serving platform, an LLM gateway for a 5,000-engineer company, or a vector search service. The core is GPU capacity, batching, the KV cache, multi-tenancy and routing. Module 2 and Module 3 cover these.

3. Add GenAI to an existing system. Examples: add AI summaries to a news feed, or semantic search to an e-commerce catalogue. The core is integration. You need to decide sync or async, how to fail gracefully when the model is slow, how to cache, and how to roll out and measure.

What interviewers look for

  • Numbers. Can you turn "10 million users" into tokens per second, GPUs and a monthly bill? Back-of-envelope math is the fastest way to stand out.
  • Trade-offs, stated explicitly. Examples: bigger model or lower latency, long context or retrieval, API or self-hosted.
  • Quality as a first-class requirement. Say how you will know the system gives good answers, before and after launch.
  • Failure handling. Cover what happens when the provider rate-limits you, the model hallucinates, or a user tries to hijack the prompt.

The next lesson turns these into a repeatable six-step structure.

Key takeaways

  • The fundamentals still apply, but the bottleneck moves from CPUs and databases to GPU memory and bandwidth.
  • LLM calls take seconds, cost cents, and return probabilistic output, so latency, cost and quality all become design dimensions.
  • Questions come in three shapes. You might build an LLM product, build LLM infrastructure, or add LLM features to an existing system.
  • Evals play the role that unit tests play in classic systems. A design without them is incomplete.

Go deeper

Finished reading? Mark it done to track your progress.