GenAI system design glossary

57 terms, each linked to the lesson that explains it properly.

Agent loop
An LLM repeatedly choosing an action (a tool call), observing the result, and deciding the next step until the task is done. Also called ReAct. Lesson: The Agent Loop
ANN (approximate nearest neighbour)
Search that finds vectors that are very likely the closest, trading a little recall for orders-of-magnitude less work than comparing against every vector. Lesson: Exact vs Approximate Nearest-Neighbour Search
BM25
The classic keyword-ranking function behind full-text search. Paired with vector search in hybrid retrieval because it catches exact terms, IDs and rare words that embeddings blur. Lesson: Hybrid Search and Reranking
BPE (byte-pair encoding)
The tokenization scheme most LLMs use. It builds a vocabulary by merging frequent byte sequences, so common words are one token and rare words split into pieces. Lesson: Tokens and Tokenizers
Chunking
Splitting documents into retrievable pieces, ideally along document structure, with the section title attached so each chunk is self-contained. Lesson: Ingestion Pipelines and Chunking
Circuit breaker
A pattern that stops sending traffic to a failing provider for a cool-down period, then probes it before restoring traffic. Lesson: Model Routing and Fallbacks
Continuous batching
Scheduling that adds new requests to the running batch at every decode step instead of waiting for the whole batch to finish. Also called in-flight batching. Lesson: Batching and Throughput
Cosine similarity
The cosine of the angle between two vectors: 1 means same direction, 0 unrelated. For unit-length vectors it equals the dot product. Lesson: Embeddings for System Designers
Cross-encoder reranker
A model that scores a query and a document together for relevance. More accurate than comparing embeddings, but too slow to run on more than the top tens or hundreds of candidates. Lesson: Hybrid Search and Reranking
Decode
The phase of generation that produces output tokens one at a time. Each step reads all model weights, so it is limited by memory bandwidth. Lesson: Why LLM Inference Is Different
Disaggregated serving
Running prefill and decode on separate GPU pools, so long prompts do not stall token streaming and each pool can be sized for its own bottleneck. Lesson: Scaling Out: Tensor, Pipeline and Expert Parallelism
Distillation
Training a small student model on a large teacher model's outputs for a specific task, for much lower cost and latency. Lesson: Distillation
Durable execution
Checkpointing a long-running workflow or agent after every step, so it resumes after crashes, deploys or approval waits instead of restarting. Lesson: State, Memory and Durable Execution
Embedding
A fixed-length vector of numbers that represents the meaning of a text, image or other item, so that similar items land close together. Lesson: Embeddings for System Designers
Expert parallelism
Placing the experts of a mixture-of-experts model on different GPUs and routing each token to the GPUs holding its chosen experts. Lesson: Scaling Out: Tensor, Pipeline and Expert Parallelism
FP8 / INT4
Reduced-precision number formats for model weights (and sometimes the KV cache). FP8 halves memory versus BF16 with little quality loss; INT4 quarters it with some. Lesson: Quantization and Speculative Decoding
Function calling (tool use)
The model returns a structured request to call a tool with arguments. Your code validates and executes it, then returns the result to the model. Lesson: Function Calling and Structured Outputs
Golden set
A curated set of inputs with expected outcomes or grading criteria, used as the regression test suite for an LLM system. Lesson: Offline Evals and Golden Sets
GraphRAG
Retrieval over a knowledge graph of entities and relationships extracted by an LLM, with community summaries for corpus-wide questions. Lesson: GraphRAG and Agentic Retrieval
Guardrails
Checks around an LLM call, such as moderation, PII redaction, topic limits, injection screening and output validation, that block or modify unsafe inputs and outputs. Lesson: Prompt Injection and Guardrails
HBM (high-bandwidth memory)
The stacked memory on data-centre GPUs. Its bandwidth, 3.35 TB/s on an H100, sets the speed limit for decode. Lesson: Why LLM Inference Is Different
HNSW
Hierarchical Navigable Small World: a layered proximity graph for ANN search. Fast and high-recall, but memory-hungry. The default index in most vector databases. Lesson: HNSW Explained
IVF (inverted file index)
An ANN index that clusters vectors around centroids and searches only the nprobe clusters nearest the query. Lesson: IVF and Product Quantization
KV cache
The stored attention keys and values for every token already processed, kept in GPU memory so each decode step does not recompute them. Grows linearly with context length and concurrent users. Lesson: GPU Memory and the KV Cache
Lethal trifecta
The combination of untrusted input, access to private data, and an external communication channel that lets prompt injection exfiltrate data. Lesson: Sandboxing and Permissions
Little's law
Requests in flight = arrival rate × time in system. Converts QPS and response time into the number of concurrent sequences a GPU fleet must hold. Lesson: Back-of-Envelope Math for LLM Systems
LLM gateway
A proxy in front of model providers that handles auth, routing, fallbacks, token-based rate limits, caching, logging and cost attribution. Lesson: The LLM Gateway Pattern
LLM-as-judge
Using a language model to grade outputs against a rubric or reference. It needs calibration against human labels and bias mitigation. Lesson: LLM-as-Judge
LoRA
Low-rank adaptation: fine-tuning by training small adapter matrices on a frozen model, typically well under 2% of its parameters. Lesson: LoRA and QLoRA
MCP (Model Context Protocol)
An open protocol for connecting AI applications to tools and data through standard servers, which turns N × M integrations into N + M. Lesson: MCP: Connecting Agents to Systems
Memory-bound vs compute-bound
Whether a workload is limited by moving bytes (decode at small batch) or by arithmetic (prefill, large batches). Decides which optimisations help. Lesson: Why LLM Inference Is Different
Mixture of experts (MoE)
A model design where each token uses only a few of many expert sub-networks. All parameters must fit in memory, but compute per token follows the active parameters. Lesson: Scaling Out: Tensor, Pipeline and Expert Parallelism
Multi-LoRA serving
Serving many LoRA adapters from one shared base model in GPU memory, batching requests for different adapters together. Lesson: Serving Many Adapters
PagedAttention
vLLM's technique of storing the KV cache in fixed-size blocks allocated on demand, like virtual memory pages, which removes most fragmentation and enables prefix sharing. Lesson: Inference Engines: vLLM, SGLang, TensorRT-LLM
pgvector
A PostgreSQL extension that adds vector columns and HNSW/IVF indexes. A strong default when vectors are in the millions and you already run Postgres. Lesson: Choosing a Vector Store
Pipeline parallelism
Splitting a model's layers into stages on different GPUs or nodes, passing activations from stage to stage. Lesson: Scaling Out: Tensor, Pipeline and Expert Parallelism
Prefill
The first phase of generation, which processes the entire prompt in one parallel pass. Compute-bound, and it determines time to first token. Lesson: Why LLM Inference Is Different
Prefix caching
Reusing the KV cache for a shared prompt prefix (system prompt, documents, chat history) across requests. Cuts prefill cost and TTFT. Exposed by APIs as prompt caching. Lesson: Inference Engines: vLLM, SGLang, TensorRT-LLM
Product quantization (PQ)
Compressing a vector by splitting it into sub-vectors and replacing each with a one-byte centroid id. Often 16 to 64 times smaller, with approximate distances. Lesson: IVF and Product Quantization
Prompt injection
Untrusted text (from a user, document, web page or tool result) that the model treats as instructions. Lesson: Prompt Injection and Guardrails
Provisioned throughput
Reserved model capacity bought at a fixed hourly price, for predictable latency and guaranteed availability at steady load. Lesson: API Costs at Scale
QLoRA
LoRA fine-tuning on a base model stored in 4-bit, which lets very large models be tuned on a single GPU. Lesson: LoRA and QLoRA
Query rewriting
Turning a conversational follow-up into a standalone search query (and filters) before retrieval. Lesson: Retrieval and Context Assembly
Recall@k
The share of the true top-k nearest neighbours that an ANN search returns. The quality metric you tune index parameters against. Lesson: Exact vs Approximate Nearest-Neighbour Search
Reciprocal rank fusion (RRF)
Merging ranked lists by summing 1 / (k + rank) for each document across lists. Needs no score calibration, which makes it the default for hybrid search. Lesson: Hybrid Search and Reranking
Roofline model
A performance model that caps throughput at the lower of peak compute and memory bandwidth × arithmetic intensity. Used to estimate LLM decode speed from GPU specs. Lesson: Capacity Planning for Inference
Semantic cache
A cache that returns a stored answer when a new question's embedding is close to a previous one. High savings, with a real risk of wrong answers. Lesson: Prompt Caching vs Semantic Caching
Server-sent events (SSE)
A one-way HTTP streaming protocol, and the standard way LLM tokens are streamed to clients. Lesson: Streaming Responses with SSE
Speculative decoding
A small draft model proposes several tokens and the large model verifies them in one step, producing multiple tokens per weight read with identical output quality. Lesson: Quantization and Speculative Decoding
Structured outputs
Constrained decoding that guarantees the model's output matches a JSON schema. Lesson: Function Calling and Structured Outputs
Temperature
A sampling setting that sharpens (low) or flattens (high) the model's next-token probability distribution. 0 always picks the top token. Lesson: Sampling, Temperature and Determinism
Tensor parallelism
Splitting each layer's matrices across GPUs that compute in parallel and all-reduce after every layer. Needs fast NVLink, so it stays within one server. Lesson: Scaling Out: Tensor, Pipeline and Expert Parallelism
Tokens per minute (TPM)
The main rate-limit unit for LLM APIs. Gateways should budget tenants in tokens, not just requests. Lesson: Rate Limiting by Tokens
TPOT (time per output token)
Average time between output tokens after the first; the inverse of per-user decode speed. Also called inter-token latency. Lesson: Why LLM Inference Is Different
Tracing
Recording each request as a tree of spans (retrieval, LLM calls, tools, guardrails) with tokens, latency and cost, to debug quality and cost issues. Lesson: Tracing and Observability
TTFT (time to first token)
Time from sending a request to receiving the first output token: queueing plus prefill. The latency users feel in a streaming UI. Lesson: Why LLM Inference Is Different