GenAI System Design
9. Fine-Tuning and Model Customization

Serving Many Adapters

How to serve hundreds of LoRA fine-tunes from one base model with multi-LoRA batching, adapter loading and caching, routing, and the capacity and isolation trade-offs.

Lesson 4 of 5 8 min

The problem

Suppose you fine-tune a model per customer, or per task: 200 LoRA adapters on an 8B base. Deploying each as its own merged model means 200 deployments, each holding a full copy of the weights on GPUs, mostly idle. At even one GPU each, that is 200 GPUs for traffic that might fit on 4.

Multi-LoRA serving

Adapters are tiny compared with the base model, so the engine keeps one base model in GPU memory and applies the right adapter per request:

Request
adapter_id = customer_42
Router / gateway
picks base-model pool
Engine batch
requests for many adapters together
Base weights
shared, loaded once
Per-request adapter math
x·W + x·B·A
  • vLLM, SGLang, TensorRT-LLM and hosted platforms support serving LoRA adapters by name per request.
  • Mixed batches: specialised kernels (from research like Punica and S-LoRA) compute the adapter contributions for a batch in which every request may use a different adapter. The base model's work is still shared across the whole batch, so batching efficiency survives.
  • Adapter overhead is small: typically a few percent extra compute per token at moderate ranks.

Adapter memory hierarchy

TierHoldsLoad time
GPU memoryHot adapters, the most recently or frequently usedInstant
CPU memoryWarm adaptersMilliseconds (PCIe transfer)
Local disk / object storageCold adapters, possibly thousandsHundreds of ms to seconds

Engines limit how many adapters can be active in one batch and how many are cached in GPU memory. Configure these from your traffic distribution: a few hot tenants and a long tail.

Routing and capacity

  • Group by base model. Each serving pool runs one base model plus many adapters. Requests for a Llama-based adapter go to the Llama pool.
  • Adapter affinity: route requests for the same adapter to the same replicas, so it stays hot in GPU memory (like prefix-cache affinity).
  • Capacity planning is by total traffic across all adapters on that base, not per adapter. That pooling is where the savings come from.
  • Noisy neighbours: one tenant's burst affects others on the pool, so apply per-tenant rate limits at the gateway (rate limiting).

Lifecycle

  • Store adapters in a registry with version, base model version, training dataset version and eval scores.
  • Deploying a new adapter version is a registry update. There's no need to redeploy the servers.
  • Base model upgrades invalidate every adapter trained on the old base. Plan retraining pipelines, which is a strong reason to keep training data and pipelines reproducible.

When to merge instead

  • A single fine-tune with high traffic: merge it into the base for zero adapter overhead and simpler serving.
  • Very high ranks or many target modules, where adapter overhead grows.
  • Engines or hardware without efficient multi-LoRA kernels.

Key takeaways

  • Multi-LoRA serving keeps one copy of the base model in GPU memory and applies a per-request adapter, so hundreds of fine-tunes share the same GPUs.
  • Engines batch requests for different adapters together (for example with S-LoRA and Punica-style kernels), so utilisation stays high.
  • Hot adapters stay in GPU memory, warm ones in CPU memory, and cold ones in object storage. Loading takes milliseconds to seconds.
  • This makes per-customer or per-task fine-tuning economical, where dedicated deployments would multiply GPU cost.

Go deeper

Finished reading? Mark it done to track your progress.