NewFree interactive course
System design for the GenAI era
Caching and sharding won't size a GPU fleet. Learn the parts of AI-native systems that interviews now ask about, including the KV cache, batching, vector indexes and token economics. Every lesson has diagrams, real numbers and calculators you can play with.
- Lessons
- 73
- Interactive calculators
- 19
- Module quizzes
- 11
- Gateway
- Orchestration
- Retrieval
- Model
- Inference
- Hardware
LLMs Explained Like System Design.
Start with foundational concepts— neural networks, tokens, embeddings, vectors, layers—and learn how they fit together without getting deep into the math. Tap to explore and learn at your own pace.
Latest Insights
Stay updated with the most important developments in AI and machine learning
FeaturedMilvus splits a vector database into stateless proxies, coordinators, specialised worker nodes and durable storage. Here is how inserts and searches flow through it, and why that design decides how it scales and fails.
FeaturedMilvus segment compaction, etcd revision compaction and etcd defrag are three different operations that get mixed up during a NOSPACE incident. Here is what each reclaims, and what to do when none of them helps.
FeaturedIn Milvus 2.5, the write stream feeds the query layer too. A bulk insert can push up search latency with flat search traffic, and the obvious fixes (more DataNodes, more replicas, weaker consistency) often miss. Here is the mechanism and what actually works.
FeaturedAI is not just automating jobs, it is creating entirely new ones, and reshaping almost every job that already exists. Here is a practical map of where the opportunities are, whether you write code or not.
FeaturedSkills are not prompts. They are reusable, invokable workflows that encode expertise once and execute reliably. Learn how to design Claude Skills with the right fields, forking strategy, tool permissions, and composition patterns.
FeaturedMCP is the open protocol that standardizes how LLMs connect to external tools and data. This deep-dive covers the architecture, primitives, transport layers, and what MCP is not, for engineers building with AI.

How to Design a Claude Skill
Skills are not prompts. They are reusable, invokable workflows that encode expertise once and execute reliably. Learn how to design Claude Skills with the right fields, forking strategy, tool permissions, and composition patterns.

What is MCP? The Model Context Protocol Explained for Engineers
MCP is the open protocol that standardizes how LLMs connect to external tools and data. This deep-dive covers the architecture, primitives, transport layers, and what MCP is not, for engineers building with AI.

What is Gemma 4? Google's Most Capable Open Model Explained
Gemma 4 is Google's open-weight model family with four sizes, 256K token context, native multimodal inputs, and Apache 2.0 licensing — and it's particularly well-suited for RAG pipelines.

MCP 2.0: What's Changing for AI Agents and RAG Pipelines
MCP 2.0 introduces long-running Task workflows, server-initiated Triggers, and native streaming — while six major AI companies formalize MCP as the universal agent standard under the Linux Foundation.

AI Digest: April 2026
Gemma 4 goes open-source, Meta drops Llama 5 with 5M token context, MCP 2.0 launches with the Agentic AI Foundation, OpenAI consolidates into a superapp, and Anthropic hits $30B run-rate.

The New SEO: Why AI Citations Are Your Brand's Next Growth Channel
Google rankings are no longer enough. ChatGPT, Perplexity, and Claude are your buyers' new search engines—and most brands are completely invisible. Here are 5 strategies to change that.

