System design for the GenAI era

You know load balancers, caches and sharding. This course covers what changes when an LLM sits in the request path: GPUs and the KV cache, batching, vector search, RAG, agents, and the numbers you need for capacity planning.

Free, no sign-up. 10 modules and 10 full design problems: 72 lessons, with a quiz at the end of every module.

The syllabus follows the stack you will be asked to design:

  1. Gateway
  2. Orchestration
  3. Retrieval
  4. Model
  5. Inference
  6. Hardware
Evals, observability, guardrails

After this course you can

  • Estimate GPUs, tokens per second and monthly cost from a DAU number, on a whiteboard, in five minutes.
  • Explain why decode is memory-bound, what the KV cache is, and how continuous batching raises throughput.
  • Choose between HNSW, IVF-PQ, pgvector and a managed vector database, and size the index.
  • Structure any GenAI design question with a six-step framework that covers quality, latency, cost and safety.

Course content

Foundations

0.How to Ace a GenAI System Design InterviewWhat changes when an LLM sits in the request path, and a six-step framework for any GenAI design question.6 lessons, 52 min, quiz
  1. How GenAI Design Rounds Differ7 min
  2. The Six-Step Framework10 min
  3. The Quality, Latency, Cost Triangle9 min
  4. Back-of-Envelope Math for LLM Systems12 min
  5. Numbers Every AI Engineer Should Know8 min
  6. Common Mistakes and How to Avoid Them6 min
  7. Module 0 quiz
1.LLM Fundamentals for System DesignersTokens, context windows, prefill and decode, and choosing between API and self-hosted models.6 lessons, 53 min, quiz
  1. Tokens and Tokenizers8 min
  2. Context Windows and Their Limits10 min
  3. Prefill vs Decode: TTFT and TPOT9 min
  4. Sampling, Temperature and Determinism9 min
  5. Model Tiers: Frontier, Mid and Small8 min
  6. API vs Self-Hosted Models9 min
  7. Module 1 quiz
2.Inference, GPUs and ServingWhy decode is memory-bound, how the KV cache eats GPU memory, and how serving engines batch requests.7 lessons, 77 min, quiz
  1. Why LLM Inference Is Different10 min
  2. GPU Memory and the KV Cache12 min
  3. Batching and Throughput11 min
  4. Inference Engines: vLLM, SGLang, TensorRT-LLM10 min
  5. Quantization and Speculative Decoding10 min
  6. Scaling Out: Tensor, Pipeline and Expert Parallelism11 min
  7. Capacity Planning for Inference13 min
  8. Module 2 quiz
3.Embeddings and Vector SearchEmbeddings, approximate nearest-neighbour indexes, and sizing a vector store for 100M documents.7 lessons, 70 min, quiz
  1. Embeddings for System Designers9 min
  2. Exact vs Approximate Nearest-Neighbour Search8 min
  3. HNSW Explained11 min
  4. IVF and Product Quantization11 min
  5. Choosing a Vector Store9 min
  6. Hybrid Search and Reranking10 min
  7. Sizing and Sharding Vector Indexes12 min
  8. Module 3 quiz
4.RAG ArchitectureIngestion, chunking, retrieval and context assembly, with freshness and permissions handled properly.7 lessons, 66 min, quiz
  1. Anatomy of a RAG System9 min
  2. Ingestion Pipelines and Chunking11 min
  3. Retrieval and Context Assembly10 min
  4. Freshness, Deletes and Re-indexing9 min
  5. Permissions and Multi-Tenancy9 min
  6. GraphRAG and Agentic Retrieval10 min
  7. RAG vs Fine-Tuning vs Long Context8 min
  8. Module 4 quiz
5.Agents and Tool UseFunction calling, the agent loop, MCP, multi-agent systems and keeping agent costs bounded.7 lessons, 63 min, quiz
  1. Function Calling and Structured Outputs9 min
  2. The Agent Loop10 min
  3. MCP: Connecting Agents to Systems8 min
  4. Multi-Agent Architectures9 min
  5. State, Memory and Durable Execution9 min
  6. Sandboxing and Permissions9 min
  7. Controlling Agent Cost and Latency9 min
  8. Module 5 quiz
6.Integrating LLMs into Existing SystemsLLM gateways, routing and fallbacks, token-based rate limits, streaming, and caching.6 lessons, 55 min, quiz
  1. The LLM Gateway Pattern9 min
  2. Model Routing and Fallbacks9 min
  3. Rate Limiting by Tokens10 min
  4. Streaming Responses with SSE9 min
  5. Batch and Async LLM Jobs8 min
  6. Prompt Caching vs Semantic Caching10 min
  7. Module 6 quiz
7.Evaluation, Observability and GuardrailsOffline and online evals, LLM-as-judge, tracing, prompt injection and PII controls.6 lessons, 55 min, quiz
  1. Why Evals Are the Test Suite8 min
  2. Offline Evals and Golden Sets10 min
  3. LLM-as-Judge9 min
  4. Tracing and Observability9 min
  5. Prompt Injection and Guardrails10 min
  6. PII, Compliance and Data Residency9 min
  7. Module 7 quiz
8.Cost and Capacity PlanningPrice per million tokens, GPU-hours, build vs buy, and a full worked estimate.5 lessons, 47 min, quiz
  1. The Unit Economics of Tokens8 min
  2. API Costs at Scale9 min
  3. Self-Hosting Costs: GPU-Hours9 min
  4. Build vs Buy9 min
  5. A Worked Estimate, End to End12 min
  6. Module 8 quiz
9.Fine-Tuning and Model CustomizationLoRA and QLoRA, training data pipelines, and serving many adapters on one base model.5 lessons, 44 min, quiz
  1. When to Fine-Tune9 min
  2. LoRA and QLoRA9 min
  3. Training Data Pipelines9 min
  4. Serving Many Adapters8 min
  5. Distillation9 min
  6. Module 9 quiz

Design Problems

Frequently asked questions

Who is this GenAI system design course for?

Software engineers who know classic system design (caching, sharding, queues) and need to design or integrate LLM-powered systems, often for an upcoming interview. No ML background is assumed.

How long does it take?

Each module takes about an hour including its quiz. Module 0 alone gives you a framework and a numbers cheat-sheet you can use in an interview the same day.

Is it free?

Yes. Every lesson, calculator and quiz is free and needs no account. Progress is saved in your browser.

How is this different from classic system design prep?

Classic prep assumes CPU-bound services with millisecond responses. LLM systems are bound by GPU memory and bandwidth, respond in seconds, cost money per token, and give non-deterministic answers. The course covers the components and trade-offs that follow from that.