System design for the GenAI era
You know load balancers, caches and sharding. This course covers what changes when an LLM sits in the request path: GPUs and the KV cache, batching, vector search, RAG, agents, and the numbers you need for capacity planning.
Free, no sign-up. 10 modules and 10 full design problems: 72 lessons, with a quiz at the end of every module.
The syllabus follows the stack you will be asked to design:
- Gatewayrouting, rate limits, caching, streaming
- Orchestrationagents, tools, MCP
- Retrievalembeddings, vector search, RAG
- Modeltokens, context, fine-tuning
- InferenceKV cache, batching, serving engines
- HardwareGPUs, memory, cost
After this course you can
- Estimate GPUs, tokens per second and monthly cost from a DAU number, on a whiteboard, in five minutes.
- Explain why decode is memory-bound, what the KV cache is, and how continuous batching raises throughput.
- Choose between HNSW, IVF-PQ, pgvector and a managed vector database, and size the index.
- Structure any GenAI design question with a six-step framework that covers quality, latency, cost and safety.
Course content
Foundations
0.How to Ace a GenAI System Design InterviewWhat changes when an LLM sits in the request path, and a six-step framework for any GenAI design question.6 lessons, 52 min, quiz
1.LLM Fundamentals for System DesignersTokens, context windows, prefill and decode, and choosing between API and self-hosted models.6 lessons, 53 min, quiz
2.Inference, GPUs and ServingWhy decode is memory-bound, how the KV cache eats GPU memory, and how serving engines batch requests.7 lessons, 77 min, quiz
- Why LLM Inference Is Different10 min
- GPU Memory and the KV Cache12 min
- Batching and Throughput11 min
- Inference Engines: vLLM, SGLang, TensorRT-LLM10 min
- Quantization and Speculative Decoding10 min
- Scaling Out: Tensor, Pipeline and Expert Parallelism11 min
- Capacity Planning for Inference13 min
- Module 2 quiz
3.Embeddings and Vector SearchEmbeddings, approximate nearest-neighbour indexes, and sizing a vector store for 100M documents.7 lessons, 70 min, quiz
4.RAG ArchitectureIngestion, chunking, retrieval and context assembly, with freshness and permissions handled properly.7 lessons, 66 min, quiz
5.Agents and Tool UseFunction calling, the agent loop, MCP, multi-agent systems and keeping agent costs bounded.7 lessons, 63 min, quiz
6.Integrating LLMs into Existing SystemsLLM gateways, routing and fallbacks, token-based rate limits, streaming, and caching.6 lessons, 55 min, quiz
7.Evaluation, Observability and GuardrailsOffline and online evals, LLM-as-judge, tracing, prompt injection and PII controls.6 lessons, 55 min, quiz
8.Cost and Capacity PlanningPrice per million tokens, GPU-hours, build vs buy, and a full worked estimate.5 lessons, 47 min, quiz
9.Fine-Tuning and Model CustomizationLoRA and QLoRA, training data pipelines, and serving many adapters on one base model.5 lessons, 44 min, quiz
Design Problems
10.GenAI System Design ProblemsTen end-to-end designs, each with a capacity estimate and an architecture diagram.10 lessons, 144 min, quiz
- Designing a ChatGPT-Style Chat Service16 min
- Designing Enterprise Document Q&A15 min
- Designing an AI Coding Assistant15 min
- Designing a Customer-Support Agent14 min
- Designing Semantic Product Search14 min
- Designing an LLM Gateway15 min
- Designing Content Moderation at Scale14 min
- Designing Text-to-SQL Analytics13 min
- Designing a Real-Time Voice Agent14 min
- Designing Embedding-Based Recommendations14 min
- Module 10 quiz
Frequently asked questions
Who is this GenAI system design course for?
Software engineers who know classic system design (caching, sharding, queues) and need to design or integrate LLM-powered systems, often for an upcoming interview. No ML background is assumed.
How long does it take?
Each module takes about an hour including its quiz. Module 0 alone gives you a framework and a numbers cheat-sheet you can use in an interview the same day.
Is it free?
Yes. Every lesson, calculator and quiz is free and needs no account. Progress is saved in your browser.
How is this different from classic system design prep?
Classic prep assumes CPU-bound services with millisecond responses. LLM systems are bound by GPU memory and bandwidth, respond in seconds, cost money per token, and give non-deterministic answers. The course covers the components and trade-offs that follow from that.