Offline Evals and Golden Sets
How to build and maintain a golden eval set, including sourcing cases, labelling, coverage, sizing it with statistics, scoring methods, and running evals as part of CI.
What goes in a golden set
Each case has:
- Input: the user message, plus conversation history or documents if relevant.
- Expected outcome: a reference answer, key facts that must appear, the correct label, the expected tool calls, or a rubric.
- Metadata: intent, difficulty, source, tags (such as
multi-hop,refusal-expected,injection).
Sourcing cases
| Source | Why |
|---|---|
| Sampled production traffic | Represents what users actually ask. Stratify by intent so rare but important intents are covered |
| Failures (thumbs-down, escalations, bug reports) | Regression protection. Every fixed bug becomes a case |
| Edge cases from domain experts | Tricky policy boundaries, ambiguous questions |
| Adversarial cases | Prompt injection, jailbreak attempts, PII bait |
| Synthetic cases generated by an LLM | Fast coverage of new features before real traffic exists. Review them, since generated cases skew easy |
Labelling: domain experts write or approve the expected outcomes. Track disagreement between labellers, because if humans disagree, the metric is noisy too.
Hygiene: remove PII or use consented data, keep a held-out subset nobody tunes prompts against (to avoid overfitting), and version the set.
How big?
How big does an eval set need to be?
A pass rate from a small eval set is a noisy estimate. See how wide the uncertainty is, and how many examples it takes to trust a small improvement.
With 100 examples, a pass rate of 82% could plausibly be anywhere from about 73% to 88%. Moving from 82% to 85% looks like progress but is well within the noise. Grow the eval set, compare variants on the same examples (paired comparison), and look at which examples flipped, not just the headline number.
Paired comparison, running both variants on the same cases and counting how many flipped each way, is much more sensitive than comparing two independent pass rates. Always look at the flipped cases.
Scoring methods
From cheapest and most reliable to most flexible:
- Exact or programmatic checks: label equals the expected label, JSON validates, a required citation is present, the SQL returns the right rows, the code passes tests. Use these wherever you can.
- Reference-based similarity: key facts present, overlap with a reference answer. It is cheap, but brittle for free-form text.
- LLM-as-judge: a model grades against a rubric or reference (next lesson). Flexible, but it needs calibration.
- Human review: the gold standard. It is slow and expensive, so use it to calibrate judges and to audit samples.
For agents, score the outcome (final state), the trajectory (were the tool calls sensible?), and the cost (steps, tokens).
Running evals as CI
- Report metrics per slice (intent, language, difficulty). An average can hide a collapse in one category.
- Track cost and latency alongside quality, to catch "better but 3× more expensive".
- Pin model versions in the baseline, so a provider-side change shows up as a diff.
- Keep a fast smoke subset (≈ 50 cases) for every commit and the full suite nightly or before release.
Key takeaways
- A golden set is a curated collection of inputs with expected outputs or grading criteria that represents real usage, edge cases and known failures.
- Source cases from production traffic, stratified by intent and difficulty, plus adversarial and regression cases.
- Size matters. 100 cases gives about ±8 points of uncertainty on a pass rate, so detecting small improvements needs hundreds to thousands.
- Prefer deterministic checks where possible, and use LLM judges for the rest.