Inference Engines: vLLM, SGLang, TensorRT-LLM
What an inference engine does beyond running the model, how PagedAttention and prefix caching work, and how to choose between vLLM, SGLang, TensorRT-LLM and hosted APIs.
What an inference engine actually does
Running a forward pass of a model is a few lines of PyTorch. Serving it to thousands of users efficiently is the engine's job:
- Scheduling: continuous batching, chunked prefill, preemption. See the batching lesson.
- Memory management: allocating, sharing and evicting KV cache.
- Kernels: fused attention (FlashAttention and variants), quantized matmuls, CUDA graphs to cut launch overhead.
- Distribution: splitting a model across GPUs (next lessons).
- Extras: structured output (JSON-schema constrained decoding), LoRA adapter serving, speculative decoding, metrics.
PagedAttention
Early servers allocated each request's KV cache as one contiguous buffer sized for the maximum possible length. Most requests finish far short of that, so most of the memory sat reserved and unused. Studies of such systems found only about 20–40% of KV memory actually holding tokens.
PagedAttention, introduced by vLLM in 2023, borrows virtual memory from operating systems:
- The KV cache is divided into fixed-size blocks, for example 16 tokens each.
- Each sequence has a block table mapping its logical positions to physical blocks anywhere in memory.
- Blocks are allocated as the sequence grows and freed when it finishes.
Waste drops to under one block per sequence, which means more concurrent sequences in the same memory, and so higher throughput. Block tables also make sharing cheap: two sequences can point at the same physical block.
Prefix caching
Many requests start with the same tokens: a 2,000-token system prompt, the same uploaded document, or a chat history that grows one turn at a time. Recomputing that prefill every time wastes compute and adds TTFT.
Prefix caching keeps the KV blocks of recent prefixes, keyed by a hash of their tokens. A new request whose prompt starts with a cached prefix skips prefill for that part. vLLM calls this automatic prefix caching. SGLang's RadixAttention organises cached prefixes in a radix tree, so partial overlaps are reused too.
This is the same mechanism behind prompt caching in hosted APIs. Cached input tokens are billed at roughly 10% of the normal price because the provider skips the prefill work.
For multi-replica deployments, prefix caching only helps if requests with the same prefix reach the same replica. The routing layer should hash on the prefix, for example on conversation ID or document ID, rather than using round-robin. This is sometimes called cache-aware or KV-aware routing.
Choosing an engine
| Option | Strengths | Trade-offs | Choose when |
|---|---|---|---|
| vLLM | Broadest model and hardware support (NVIDIA, AMD, TPU, and others), large community, OpenAI-compatible server | Not always the fastest on a given setup | The default for self-hosting open models |
| SGLang | RadixAttention prefix reuse, fast structured output, strong on agentic and multi-turn workloads | Smaller ecosystem than vLLM | Heavy prefix sharing, many short structured calls |
| TensorRT-LLM (often behind NVIDIA Triton or Dynamo) | Peak performance on NVIDIA through compiled engines, FP8 and FP4 kernels | NVIDIA-only, build step per model and GPU, less flexible | Very high, steady volume on a fixed model |
| llama.cpp / Ollama | Runs on CPUs, Macs, small GPUs, GGUF quantization | Not built for high-concurrency serving | Local, edge, dev environments |
| Hosted API | Zero ops, frontier models, elastic | Per-token cost, rate limits, data leaves your boundary | Most products, especially below a few hundred QPS |
For deployment, engines run as containers behind a load balancer. Kubernetes projects such as KServe, llm-d and NVIDIA Dynamo add model-aware routing, autoscaling on queue depth or KV cache usage, and disaggregated prefill/decode.
Key takeaways
- An inference engine handles scheduling, batching, KV cache memory management and optimised kernels. The model math is the easy part.
- PagedAttention stores the KV cache in fixed-size blocks allocated on demand, which nearly eliminates fragmentation.
- Prefix caching reuses KV blocks for shared prompt prefixes, cutting TTFT and cost for system prompts, documents and chat history.
- vLLM and SGLang are the open defaults. TensorRT-LLM trades flexibility for peak NVIDIA performance.