GenAI System Design
9. Fine-Tuning and Model Customization

LoRA and QLoRA

How low-rank adaptation fine-tunes a model by training small adapter matrices, why it cuts training memory dramatically, how QLoRA adds 4-bit quantization, and the key hyperparameters.

Lesson 2 of 5 9 min

Why full fine-tuning is expensive

Full fine-tuning updates every weight. With the Adam optimizer in mixed precision, each parameter needs roughly:

  • 2 bytes: BF16 weight
  • 2 bytes: gradient
  • 12 bytes: FP32 master weight plus two Adam moments

That is ≈ 16 bytes per parameter before activations. An 8B model needs ≈ 128 GB and a 70B model ≈ 1.1 TB, which means multi-GPU or multi-node training with sharding (FSDP, DeepSpeed ZeRO).

LoRA: low-rank adaptation

Instead of updating a big weight matrix W (d × d), LoRA freezes W and learns a small update ΔW = B·A, where A is r × d and B is d × r, with rank r much smaller than d (for example 8–64).

Input x
Frozen W
d × d, not trained
+ B·A
rank r adapter, trained
Output
Wx + BAx

For d = 4,096 and r = 16, the adapter has 2 × 4,096 × 16 ≈ 131K parameters, versus 16.8M for the full matrix, which is 128× fewer. Applied across attention and MLP layers, a LoRA typically trains 0.1–2% of the model's parameters.

Consequences:

  • Memory: gradients and optimizer state exist only for the adapter. You need the frozen weights (2 bytes per parameter in BF16) plus a small overhead.
  • Speed: training is faster, and you can use fewer or smaller GPUs.
  • Storage: an adapter is a small file of megabytes to a few hundred MB, so you can keep hundreds.
  • Quality: for most behaviour-shaping tasks, LoRA gets close to full fine-tuning. For big shifts, such as a new language or heavy domain adaptation, full fine-tuning can still win.

QLoRA

QLoRA quantizes the frozen base model to 4-bit (the NF4 format) and trains LoRA adapters on top in higher precision. The frozen weights then cost ≈ 0.5 bytes per parameter:

Training memory: full vs LoRA vs QLoRA

Full fine-tuning stores gradients and optimizer state for every weight. LoRA trains a tiny adapter on a frozen model. QLoRA also shrinks the frozen model to 4 bits.

Full fine-tuning131 GB, 2 × H100 80 GB
BF16 weights + gradients + FP32 Adam states for every parameter
LoRA22 GB, 1 × H100 80 GB
Frozen BF16 base, small trainable adapters (≈ 1% of params)
QLoRA11 GB, 1 × H100 80 GB
Frozen 4-bit base, adapters trained in higher precision
WeightsGradients + optimizerActivationsOverhead

Full fine-tuning needs roughly 16 bytes per parameter before activations, about 8× the model's inference weights. LoRA removes most of that, and QLoRA lets a 70B model train on a single 80 GB GPU. These are estimates with gradient checkpointing. Real usage depends on the framework and on sharding (FSDP, DeepSpeed).

This made fine-tuning 30B–70B models possible on a single GPU. The costs are somewhat slower training (weights are dequantized on the fly) and small quality differences compared with BF16 LoRA.

Key hyperparameters

ParameterWhat it controlsTypical
Rank rAdapter capacity8–64. Higher for harder tasks
AlphaScaling of the updateOften r or 2r
Target modulesWhich layers get adaptersAttention projections, and usually the MLP layers too
Learning rateStep sizeHigher than full fine-tuning, often around 1e-4 to 2e-4
EpochsPasses over the data1–3. More risks overfitting on small datasets

Tune them against your eval set, and watch for overfitting: training loss keeps falling while eval quality stalls or drops.

Merge or keep separate?

  • Merge the adapter into the base weights (W + BA) to get a normal model with no inference overhead. This suits a single dedicated deployment.
  • Keep it separate to serve many adapters on one base model, switching per request (serving many adapters). This suits multi-tenant or multi-task platforms.

Key takeaways

  • LoRA freezes the base model and trains small low-rank matrices added to chosen layers, typically well under 1% of parameters.
  • Training memory drops from about 16 bytes per parameter (full fine-tuning with Adam) to roughly the frozen weights plus a small adapter overhead.
  • QLoRA stores the frozen base in 4-bit, so a 70B model can be fine-tuned on a single 80 GB GPU.
  • Adapters are small files (MB to hundreds of MB). They can be swapped at serve time, or merged into the base weights.

Go deeper

Finished reading? Mark it done to track your progress.