GenAI System Design
9. Fine-Tuning and Model Customization

Distillation

How to transfer a large model's behaviour on a specific task into a small, cheap, fast model, covering teacher-generated data, training, evaluation, the economics, and the risks.

Lesson 5 of 5 9 min

The idea

Frontier models are great generalists but expensive and slow. Most production tasks are narrow: classify this ticket, extract these fields, write this kind of summary. A small model (1–8B parameters) trained specifically on the big model's outputs for that task can often come close to the big model's quality on that task, at a fraction of the cost and latency.

Real inputs
production traffic
Teacher
frontier model, best prompt
Filter
graders, rules, human spot-checks
Student fine-tune
small model, LoRA or full
Eval vs teacher
held-out set

Recipe

  1. Collect inputs that represent real traffic, tens of thousands for complex tasks, stratified by intent and difficulty.
  2. Generate teacher outputs with your best prompt, tools and settings. Optionally include the teacher's reasoning. Training students on step-by-step rationales often improves them on reasoning-heavy tasks.
  3. Filter: keep only outputs that pass validation, graders or human checks. The student learns everything you show it, mistakes included.
  4. Fine-tune the student (SFT, often LoRA). Optionally follow with preference tuning, using teacher-ranked pairs.
  5. Evaluate on a held-out set against the teacher and against the business metric, not just similarity to teacher outputs.
  6. Deploy with a fallback: route low-confidence or out-of-scope inputs to the teacher.

The economics

Example: a classification and extraction step at 20M requests a month, 1K input and 100 output tokens each.

  • Frontier model at $5 / $25 per million: 20B input × $5 + 2B output × $25 = $150K a month
  • Distilled 8B model self-hosted: a few GPUs with high batching. Even at 10 H100s that is ≈ $18K a month, with much lower latency
  • One-off cost: generating teacher data (for example 100K examples ≈ 110M tokens, under $1,000 on a frontier model) plus training (GPU-hours) plus engineering

At this volume the payback period is days to weeks. At low volume, distillation isn't worth the engineering. Use the frontier model with good prompts.

Distillation vs quantization vs routing

TechniqueWhat it doesKeeps generality?
QuantizationSame model, fewer bits per weightYes (small quality loss)
RoutingSend easy requests to an existing small modelYes, per request
DistillationTrain a small model to specialiseNo. It's a task specialist

These combine: distil into an 8B model, quantize it to FP8 or INT4 for serving, and route anything it can't handle to the frontier model.

Risks

  • Distribution shift: the student is only good where it was trained. New product areas, languages or input formats need new data. Monitor online quality, and route or retrain.
  • Inherited errors: the teacher's systematic mistakes become the student's. Filtering and human review mitigate this.
  • Lost safety behaviours: a narrowly trained student may lose some refusal and safety properties. Include safety cases in training and evals.
  • Licensing and terms: many providers restrict using their outputs to train competing models. Check the terms, or use open-weight teachers whose licences allow it.
  • Maintenance: when the teacher or task changes, regenerate data and retrain. Keep the pipeline automated.

Key takeaways

  • Distillation trains a small student model to imitate a large teacher's outputs on your task, for large cost and latency savings.
  • The usual recipe is to collect real inputs, generate teacher outputs (often with reasoning), filter them, fine-tune the student, and evaluate against the teacher.
  • Distilled models are specialists. They excel on the target distribution and degrade outside it, so route out-of-scope traffic to the big model.
  • Check provider terms before using a model's outputs to train another model.

Go deeper

Finished reading? Mark it done to track your progress.