GenAI System Design
9. Fine-Tuning and Model Customization

When to Fine-Tune

What fine-tuning is good and bad at, the kinds of fine-tuning (supervised, preference, reinforcement), a decision ladder from prompting to training, and what it takes to do well.

Lesson 1 of 5 9 min

What fine-tuning does

A pretrained model is further trained on your examples, so its weights shift toward the behaviour those examples show. It is good at teaching how to respond:

  • Format and structure: always produce this schema, this report layout, this citation style.
  • Style and voice: your brand tone, concise clinical notes, legal drafting conventions.
  • Narrow tasks: classify into your 200 categories, extract fields from your document types, route tickets.
  • Domain language: jargon, abbreviations and conventions the base model handles poorly.
  • Efficiency: make a small model do one task as well as a large one (distillation, lesson 5).

It is bad at teaching what is true: specific, changing facts. For that, use retrieval (see RAG vs fine-tuning vs long context).

Kinds of fine-tuning

TypeDataTeaches
Supervised fine-tuning (SFT)Input → ideal output pairsImitate the target outputs. The most common type
Preference tuning (DPO and similar)Input + preferred and rejected outputsPrefer better responses, for tone, safety and helpfulness
Reinforcement fine-tuningInputs + a grader or reward functionOptimise for a verifiable outcome, such as correct answers or passing tests
Continued pretrainingLarge unlabelled domain textDeep domain adaptation. Expensive and rarely needed

Each can be done as full fine-tuning (update all weights) or parameter-efficient fine-tuning such as LoRA (next lesson).

The decision ladder

1. Prompting
clear instructions, format spec
2. Few-shot examples
in the prompt
3. RAG / tools
missing knowledge
4. Fine-tune
behaviour gap remains
5. Distil
cut cost at scale

Move up a rung only when your eval set shows the current rung isn't good enough. Each rung costs more to build and maintain.

Good reasons to fine-tune

  • The prompt is huge and repeated: thousands of tokens of instructions and examples on every call. Training them in makes calls cheaper and faster.
  • You need a smaller, cheaper or faster model for a high-volume task, and a frontier model's outputs can serve as training data.
  • Consistency matters and prompting still gives format drift or tone variance on 5% of cases.
  • A specialised task has plenty of labelled data, such as classification or extraction, where fine-tuned small models often beat much larger prompted ones.
  • Self-hosted or on-device deployment needs a small model that does one job well.

Bad reasons

  • "So the model knows our documents." Use RAG.
  • "Our prompt engineering is messy." Fix the prompt and build evals first.
  • "Fine-tuning is what serious AI teams do." It's a cost and maintenance burden that needs to be justified.

What it takes

  • Data: hundreds to thousands of high-quality examples for SFT, more for harder tasks (data pipelines).
  • Evals: a held-out set to show the tuned model beats the baseline, and doesn't regress on general ability or safety.
  • Infrastructure: hosted fine-tuning APIs (simplest), or your own training on GPUs with LoRA or full fine-tuning.
  • Serving: a custom model to host, or adapters to serve (serving many adapters).
  • Maintenance: when a better base model ships, you may need to redo it. Keep the data and pipeline reproducible.

Key takeaways

  • Fine-tuning changes model weights to change behaviour, such as format, style, domain skills or narrow task accuracy. It is a poor way to add facts.
  • Climb the ladder. Try prompting, then few-shot examples, then RAG, and fine-tune only when evals show a gap those can't close.
  • The strongest reasons to fine-tune are consistent behaviour at scale, moving to a cheaper or faster model, and specialised tasks with many examples.
  • Fine-tuning brings ongoing costs, including data curation, eval, hosting a custom model, and redoing it when base models improve.

Go deeper

Finished reading? Mark it done to track your progress.