Training Data Pipelines
How to source, clean, label and version fine-tuning data, including using production logs, synthetic data from stronger models, deduplication, decontamination, and privacy.
Lesson 3 of 5 9 min
The pipeline
Source
logs, experts, synthetic
Filter + clean
PII, quality, dedup
Label / verify
humans or graders
Split
train / val / held-out eval
Version + train
Sources
| Source | Pros | Cons |
|---|---|---|
| Production traffic + good outputs | Matches real inputs exactly | Needs consent and privacy handling. Outputs may be mediocre |
| Human-corrected outputs (such as the edited drafts agents actually sent) | High quality, and captures the target behaviour | Slow. Depends on the workflow capturing edits |
| Expert-written examples | Highest quality for hard cases | Expensive, small |
| Synthetic from a stronger model | Fast, scalable, cheap | Inherits the teacher's errors and style. Check provider terms on training use |
| Public datasets | Free | Rarely match your task. Licence restrictions |
A common, effective recipe: real inputs from production, outputs from a strong model or corrected by humans, verified by programmatic checks or a judge, then spot-checked by experts.
Cleaning
- Deduplicate: exact and near-duplicate inputs cause overfitting to repeated cases.
- Filter quality: drop outputs that fail schema checks, judges or business rules. One bad example can teach a bad habit.
- Balance: make sure rare but important categories are represented. Don't let 80% of examples be the easiest intent.
- Consistency: examples must agree with each other. If two labellers format dates differently, the model learns both.
- Decontaminate: remove any examples that overlap your eval set. Otherwise eval scores are inflated and meaningless.
- Format: match exactly how the model will be prompted in production, including the system prompt, chat template and tool schemas.
How much data?
Plot a learning curve: train on 25%, 50% and 100% of the data and measure eval quality. If quality is still climbing, more data helps. If it's flat, work on quality or task design instead.
Privacy and rights
- Only train on user data you have the right and consent to use. Check your terms of service and customer contracts.
- Scrub PII before training. Models can memorise and regurgitate training examples.
- Remember you can't delete one person's data from trained weights. Deletion means retraining without it. Prefer RAG for per-user or per-customer knowledge.
- For multi-tenant products, never train a shared model on one customer's confidential data without explicit agreement. Per-tenant adapters trained only on that tenant's data are the safer pattern.
Versioning and reproducibility
- Version datasets alongside the code and config that trained each model (with dataset hashes).
- Record which dataset version, base model version and hyperparameters produced each checkpoint.
- Keep the held-out eval set frozen across runs, so results are comparable.
- Automate the pipeline, so rebuilding on a new base model is a re-run, not a project.
Key takeaways
- Data quality beats quantity. A few thousand clean, consistent, representative examples usually beat a large, noisy set.
- Sources include curated production traffic, human-written examples, and synthetic outputs from stronger models, each with trade-offs.
- Clean aggressively. Deduplicate, filter low quality, balance classes, and remove anything that overlaps your eval set.
- Version datasets like code, and keep personal data out unless you have consent and can handle deletion.
Go deeper
Finished reading? Mark it done to track your progress.