GenAI System Design
9. Fine-Tuning and Model Customization

Training Data Pipelines

How to source, clean, label and version fine-tuning data, including using production logs, synthetic data from stronger models, deduplication, decontamination, and privacy.

Lesson 3 of 5 9 min

The pipeline

Source
logs, experts, synthetic
Filter + clean
PII, quality, dedup
Label / verify
humans or graders
Split
train / val / held-out eval
Version + train

Sources

SourceProsCons
Production traffic + good outputsMatches real inputs exactlyNeeds consent and privacy handling. Outputs may be mediocre
Human-corrected outputs (such as the edited drafts agents actually sent)High quality, and captures the target behaviourSlow. Depends on the workflow capturing edits
Expert-written examplesHighest quality for hard casesExpensive, small
Synthetic from a stronger modelFast, scalable, cheapInherits the teacher's errors and style. Check provider terms on training use
Public datasetsFreeRarely match your task. Licence restrictions

A common, effective recipe: real inputs from production, outputs from a strong model or corrected by humans, verified by programmatic checks or a judge, then spot-checked by experts.

Cleaning

  • Deduplicate: exact and near-duplicate inputs cause overfitting to repeated cases.
  • Filter quality: drop outputs that fail schema checks, judges or business rules. One bad example can teach a bad habit.
  • Balance: make sure rare but important categories are represented. Don't let 80% of examples be the easiest intent.
  • Consistency: examples must agree with each other. If two labellers format dates differently, the model learns both.
  • Decontaminate: remove any examples that overlap your eval set. Otherwise eval scores are inflated and meaningless.
  • Format: match exactly how the model will be prompted in production, including the system prompt, chat template and tool schemas.

How much data?

Plot a learning curve: train on 25%, 50% and 100% of the data and measure eval quality. If quality is still climbing, more data helps. If it's flat, work on quality or task design instead.

Privacy and rights

  • Only train on user data you have the right and consent to use. Check your terms of service and customer contracts.
  • Scrub PII before training. Models can memorise and regurgitate training examples.
  • Remember you can't delete one person's data from trained weights. Deletion means retraining without it. Prefer RAG for per-user or per-customer knowledge.
  • For multi-tenant products, never train a shared model on one customer's confidential data without explicit agreement. Per-tenant adapters trained only on that tenant's data are the safer pattern.

Versioning and reproducibility

  • Version datasets alongside the code and config that trained each model (with dataset hashes).
  • Record which dataset version, base model version and hyperparameters produced each checkpoint.
  • Keep the held-out eval set frozen across runs, so results are comparable.
  • Automate the pipeline, so rebuilding on a new base model is a re-run, not a project.

Key takeaways

  • Data quality beats quantity. A few thousand clean, consistent, representative examples usually beat a large, noisy set.
  • Sources include curated production traffic, human-written examples, and synthetic outputs from stronger models, each with trade-offs.
  • Clean aggressively. Deduplicate, filter low quality, balance classes, and remove anything that overlaps your eval set.
  • Version datasets like code, and keep personal data out unless you have consent and can handle deletion.

Go deeper

Finished reading? Mark it done to track your progress.