Serving Many Adapters
How to serve hundreds of LoRA fine-tunes from one base model with multi-LoRA batching, adapter loading and caching, routing, and the capacity and isolation trade-offs.
The problem
Suppose you fine-tune a model per customer, or per task: 200 LoRA adapters on an 8B base. Deploying each as its own merged model means 200 deployments, each holding a full copy of the weights on GPUs, mostly idle. At even one GPU each, that is 200 GPUs for traffic that might fit on 4.
Multi-LoRA serving
Adapters are tiny compared with the base model, so the engine keeps one base model in GPU memory and applies the right adapter per request:
- vLLM, SGLang, TensorRT-LLM and hosted platforms support serving LoRA adapters by name per request.
- Mixed batches: specialised kernels (from research like Punica and S-LoRA) compute the adapter contributions for a batch in which every request may use a different adapter. The base model's work is still shared across the whole batch, so batching efficiency survives.
- Adapter overhead is small: typically a few percent extra compute per token at moderate ranks.
Adapter memory hierarchy
| Tier | Holds | Load time |
|---|---|---|
| GPU memory | Hot adapters, the most recently or frequently used | Instant |
| CPU memory | Warm adapters | Milliseconds (PCIe transfer) |
| Local disk / object storage | Cold adapters, possibly thousands | Hundreds of ms to seconds |
Engines limit how many adapters can be active in one batch and how many are cached in GPU memory. Configure these from your traffic distribution: a few hot tenants and a long tail.
Routing and capacity
- Group by base model. Each serving pool runs one base model plus many adapters. Requests for a Llama-based adapter go to the Llama pool.
- Adapter affinity: route requests for the same adapter to the same replicas, so it stays hot in GPU memory (like prefix-cache affinity).
- Capacity planning is by total traffic across all adapters on that base, not per adapter. That pooling is where the savings come from.
- Noisy neighbours: one tenant's burst affects others on the pool, so apply per-tenant rate limits at the gateway (rate limiting).
Lifecycle
- Store adapters in a registry with version, base model version, training dataset version and eval scores.
- Deploying a new adapter version is a registry update. There's no need to redeploy the servers.
- Base model upgrades invalidate every adapter trained on the old base. Plan retraining pipelines, which is a strong reason to keep training data and pipelines reproducible.
When to merge instead
- A single fine-tune with high traffic: merge it into the base for zero adapter overhead and simpler serving.
- Very high ranks or many target modules, where adapter overhead grows.
- Engines or hardware without efficient multi-LoRA kernels.
Key takeaways
- Multi-LoRA serving keeps one copy of the base model in GPU memory and applies a per-request adapter, so hundreds of fine-tunes share the same GPUs.
- Engines batch requests for different adapters together (for example with S-LoRA and Punica-style kernels), so utilisation stays high.
- Hot adapters stay in GPU memory, warm ones in CPU memory, and cold ones in object storage. Loading takes milliseconds to seconds.
- This makes per-customer or per-task fine-tuning economical, where dedicated deployments would multiply GPU cost.