LoRA and QLoRA
How low-rank adaptation fine-tunes a model by training small adapter matrices, why it cuts training memory dramatically, how QLoRA adds 4-bit quantization, and the key hyperparameters.
Why full fine-tuning is expensive
Full fine-tuning updates every weight. With the Adam optimizer in mixed precision, each parameter needs roughly:
- 2 bytes: BF16 weight
- 2 bytes: gradient
- 12 bytes: FP32 master weight plus two Adam moments
That is ≈ 16 bytes per parameter before activations. An 8B model needs ≈ 128 GB and a 70B model ≈ 1.1 TB, which means multi-GPU or multi-node training with sharding (FSDP, DeepSpeed ZeRO).
LoRA: low-rank adaptation
Instead of updating a big weight matrix W (d × d), LoRA freezes W and learns a small update ΔW = B·A, where A is r × d and B is d × r, with rank r much smaller than d (for example 8–64).
For d = 4,096 and r = 16, the adapter has 2 × 4,096 × 16 ≈ 131K parameters, versus 16.8M for the full matrix, which is 128× fewer. Applied across attention and MLP layers, a LoRA typically trains 0.1–2% of the model's parameters.
Consequences:
- Memory: gradients and optimizer state exist only for the adapter. You need the frozen weights (2 bytes per parameter in BF16) plus a small overhead.
- Speed: training is faster, and you can use fewer or smaller GPUs.
- Storage: an adapter is a small file of megabytes to a few hundred MB, so you can keep hundreds.
- Quality: for most behaviour-shaping tasks, LoRA gets close to full fine-tuning. For big shifts, such as a new language or heavy domain adaptation, full fine-tuning can still win.
QLoRA
QLoRA quantizes the frozen base model to 4-bit (the NF4 format) and trains LoRA adapters on top in higher precision. The frozen weights then cost ≈ 0.5 bytes per parameter:
Training memory: full vs LoRA vs QLoRA
Full fine-tuning stores gradients and optimizer state for every weight. LoRA trains a tiny adapter on a frozen model. QLoRA also shrinks the frozen model to 4 bits.
Full fine-tuning needs roughly 16 bytes per parameter before activations, about 8× the model's inference weights. LoRA removes most of that, and QLoRA lets a 70B model train on a single 80 GB GPU. These are estimates with gradient checkpointing. Real usage depends on the framework and on sharding (FSDP, DeepSpeed).
This made fine-tuning 30B–70B models possible on a single GPU. The costs are somewhat slower training (weights are dequantized on the fly) and small quality differences compared with BF16 LoRA.
Key hyperparameters
| Parameter | What it controls | Typical |
|---|---|---|
| Rank r | Adapter capacity | 8–64. Higher for harder tasks |
| Alpha | Scaling of the update | Often r or 2r |
| Target modules | Which layers get adapters | Attention projections, and usually the MLP layers too |
| Learning rate | Step size | Higher than full fine-tuning, often around 1e-4 to 2e-4 |
| Epochs | Passes over the data | 1–3. More risks overfitting on small datasets |
Tune them against your eval set, and watch for overfitting: training loss keeps falling while eval quality stalls or drops.
Merge or keep separate?
- Merge the adapter into the base weights (W + BA) to get a normal model with no inference overhead. This suits a single dedicated deployment.
- Keep it separate to serve many adapters on one base model, switching per request (serving many adapters). This suits multi-tenant or multi-task platforms.
Key takeaways
- LoRA freezes the base model and trains small low-rank matrices added to chosen layers, typically well under 1% of parameters.
- Training memory drops from about 16 bytes per parameter (full fine-tuning with Adam) to roughly the frozen weights plus a small adapter overhead.
- QLoRA stores the frozen base in 4-bit, so a 70B model can be fine-tuned on a single 80 GB GPU.
- Adapters are small files (MB to hundreds of MB). They can be swapped at serve time, or merged into the base weights.