GenAI System Design
2. Inference, GPUs and Serving

Inference Engines: vLLM, SGLang, TensorRT-LLM

What an inference engine does beyond running the model, how PagedAttention and prefix caching work, and how to choose between vLLM, SGLang, TensorRT-LLM and hosted APIs.

Lesson 4 of 7 10 min

What an inference engine actually does

Running a forward pass of a model is a few lines of PyTorch. Serving it to thousands of users efficiently is the engine's job:

API server
OpenAI-compatible HTTP, streaming
Scheduler
continuous batching, priorities
KV cache manager
paged blocks, prefix cache
Model executor
fused kernels, CUDA graphs
GPUs
tensor / pipeline parallel
  • Scheduling: continuous batching, chunked prefill, preemption. See the batching lesson.
  • Memory management: allocating, sharing and evicting KV cache.
  • Kernels: fused attention (FlashAttention and variants), quantized matmuls, CUDA graphs to cut launch overhead.
  • Distribution: splitting a model across GPUs (next lessons).
  • Extras: structured output (JSON-schema constrained decoding), LoRA adapter serving, speculative decoding, metrics.

PagedAttention

Early servers allocated each request's KV cache as one contiguous buffer sized for the maximum possible length. Most requests finish far short of that, so most of the memory sat reserved and unused. Studies of such systems found only about 20–40% of KV memory actually holding tokens.

Pre-allocated per requestReq 1Req 2Req 3Dashed = reserved for max length but unused.Here 24 of 48 slots sit idle.PagedAttention: 4-token blocks on demandBlocks need not be contiguous; a block table maps eachsequence to its pages. Free blocks serve new requests.
Reserving max-length buffers per request wastes most of the KV cache. Fixed-size pages allocated on demand keep waste under one block per sequence.

PagedAttention, introduced by vLLM in 2023, borrows virtual memory from operating systems:

  • The KV cache is divided into fixed-size blocks, for example 16 tokens each.
  • Each sequence has a block table mapping its logical positions to physical blocks anywhere in memory.
  • Blocks are allocated as the sequence grows and freed when it finishes.

Waste drops to under one block per sequence, which means more concurrent sequences in the same memory, and so higher throughput. Block tables also make sharing cheap: two sequences can point at the same physical block.

Prefix caching

Many requests start with the same tokens: a 2,000-token system prompt, the same uploaded document, or a chat history that grows one turn at a time. Recomputing that prefill every time wastes compute and adds TTFT.

Prefix caching keeps the KV blocks of recent prefixes, keyed by a hash of their tokens. A new request whose prompt starts with a cached prefix skips prefill for that part. vLLM calls this automatic prefix caching. SGLang's RadixAttention organises cached prefixes in a radix tree, so partial overlaps are reused too.

This is the same mechanism behind prompt caching in hosted APIs. Cached input tokens are billed at roughly 10% of the normal price because the provider skips the prefill work.

For multi-replica deployments, prefix caching only helps if requests with the same prefix reach the same replica. The routing layer should hash on the prefix, for example on conversation ID or document ID, rather than using round-robin. This is sometimes called cache-aware or KV-aware routing.

Choosing an engine

OptionStrengthsTrade-offsChoose when
vLLMBroadest model and hardware support (NVIDIA, AMD, TPU, and others), large community, OpenAI-compatible serverNot always the fastest on a given setupThe default for self-hosting open models
SGLangRadixAttention prefix reuse, fast structured output, strong on agentic and multi-turn workloadsSmaller ecosystem than vLLMHeavy prefix sharing, many short structured calls
TensorRT-LLM (often behind NVIDIA Triton or Dynamo)Peak performance on NVIDIA through compiled engines, FP8 and FP4 kernelsNVIDIA-only, build step per model and GPU, less flexibleVery high, steady volume on a fixed model
llama.cpp / OllamaRuns on CPUs, Macs, small GPUs, GGUF quantizationNot built for high-concurrency servingLocal, edge, dev environments
Hosted APIZero ops, frontier models, elasticPer-token cost, rate limits, data leaves your boundaryMost products, especially below a few hundred QPS

For deployment, engines run as containers behind a load balancer. Kubernetes projects such as KServe, llm-d and NVIDIA Dynamo add model-aware routing, autoscaling on queue depth or KV cache usage, and disaggregated prefill/decode.

Key takeaways

  • An inference engine handles scheduling, batching, KV cache memory management and optimised kernels. The model math is the easy part.
  • PagedAttention stores the KV cache in fixed-size blocks allocated on demand, which nearly eliminates fragmentation.
  • Prefix caching reuses KV blocks for shared prompt prefixes, cutting TTFT and cost for system prompts, documents and chat history.
  • vLLM and SGLang are the open defaults. TensorRT-LLM trades flexibility for peak NVIDIA performance.

Go deeper

Finished reading? Mark it done to track your progress.