When to Fine-Tune
What fine-tuning is good and bad at, the kinds of fine-tuning (supervised, preference, reinforcement), a decision ladder from prompting to training, and what it takes to do well.
Lesson 1 of 5 9 min
What fine-tuning does
A pretrained model is further trained on your examples, so its weights shift toward the behaviour those examples show. It is good at teaching how to respond:
- Format and structure: always produce this schema, this report layout, this citation style.
- Style and voice: your brand tone, concise clinical notes, legal drafting conventions.
- Narrow tasks: classify into your 200 categories, extract fields from your document types, route tickets.
- Domain language: jargon, abbreviations and conventions the base model handles poorly.
- Efficiency: make a small model do one task as well as a large one (distillation, lesson 5).
It is bad at teaching what is true: specific, changing facts. For that, use retrieval (see RAG vs fine-tuning vs long context).
Kinds of fine-tuning
| Type | Data | Teaches |
|---|---|---|
| Supervised fine-tuning (SFT) | Input → ideal output pairs | Imitate the target outputs. The most common type |
| Preference tuning (DPO and similar) | Input + preferred and rejected outputs | Prefer better responses, for tone, safety and helpfulness |
| Reinforcement fine-tuning | Inputs + a grader or reward function | Optimise for a verifiable outcome, such as correct answers or passing tests |
| Continued pretraining | Large unlabelled domain text | Deep domain adaptation. Expensive and rarely needed |
Each can be done as full fine-tuning (update all weights) or parameter-efficient fine-tuning such as LoRA (next lesson).
The decision ladder
1. Prompting
clear instructions, format spec
2. Few-shot examples
in the prompt
3. RAG / tools
missing knowledge
4. Fine-tune
behaviour gap remains
5. Distil
cut cost at scale
Move up a rung only when your eval set shows the current rung isn't good enough. Each rung costs more to build and maintain.
Good reasons to fine-tune
- The prompt is huge and repeated: thousands of tokens of instructions and examples on every call. Training them in makes calls cheaper and faster.
- You need a smaller, cheaper or faster model for a high-volume task, and a frontier model's outputs can serve as training data.
- Consistency matters and prompting still gives format drift or tone variance on 5% of cases.
- A specialised task has plenty of labelled data, such as classification or extraction, where fine-tuned small models often beat much larger prompted ones.
- Self-hosted or on-device deployment needs a small model that does one job well.
Bad reasons
- "So the model knows our documents." Use RAG.
- "Our prompt engineering is messy." Fix the prompt and build evals first.
- "Fine-tuning is what serious AI teams do." It's a cost and maintenance burden that needs to be justified.
What it takes
- Data: hundreds to thousands of high-quality examples for SFT, more for harder tasks (data pipelines).
- Evals: a held-out set to show the tuned model beats the baseline, and doesn't regress on general ability or safety.
- Infrastructure: hosted fine-tuning APIs (simplest), or your own training on GPUs with LoRA or full fine-tuning.
- Serving: a custom model to host, or adapters to serve (serving many adapters).
- Maintenance: when a better base model ships, you may need to redo it. Keep the data and pipeline reproducible.
Key takeaways
- Fine-tuning changes model weights to change behaviour, such as format, style, domain skills or narrow task accuracy. It is a poor way to add facts.
- Climb the ladder. Try prompting, then few-shot examples, then RAG, and fine-tune only when evals show a gap those can't close.
- The strongest reasons to fine-tune are consistent behaviour at scale, moving to a cheaper or faster model, and specialised tasks with many examples.
- Fine-tuning brings ongoing costs, including data curation, eval, hosting a custom model, and redoing it when base models improve.
Go deeper
Finished reading? Mark it done to track your progress.