LLM-as-Judge
Using a language model to grade outputs at scale, covering rubric design, pointwise vs pairwise judging, known biases, calibrating against human labels, and cost.
Why use a model to grade
Many qualities can't be checked with code: Is this answer faithful to the sources? Does it address the question? Is the tone right? Human review is accurate but costs minutes and dollars per case. An LLM judge reads the input, the output (and optionally a reference or sources), and returns a score with a rationale, in seconds and for cents or less.
Judge setups
| Setup | Question it answers | Use for |
|---|---|---|
| Pointwise with rubric | "Does this answer meet criteria X, Y, Z?" | Absolute quality tracking, CI gates |
| Reference-based | "Does this answer contain the facts in the reference?" | Factual Q&A with known answers |
| Pairwise | "Which of A and B is better?" | Comparing prompt or model variants. More sensitive than pointwise |
| Groundedness / faithfulness | "Is every claim supported by these sources?" | RAG hallucination detection |
Writing a good rubric
Vague: "Rate the answer's quality from 1 to 10." The scores are noisy and cluster around 7–8.
Better: break quality into specific, observable criteria, scored on a small scale or pass/fail:
Given the QUESTION, the SOURCES and the ANSWER, evaluate:
1. Correctness: Does the answer state the refund window accurately per the sources? (pass/fail)
2. Groundedness: Is every factual claim supported by a cited source? List any unsupported claims. (pass/fail)
3. Completeness: Does it address all parts of the question? (0 = no, 1 = partially, 2 = fully)
First explain your reasoning briefly, then output JSON: {"correct": bool, "grounded": bool, "complete": 0|1|2}- Ask for reasoning before the verdict. It improves consistency.
- Use structured output so scores are machine-readable.
- Include examples of passing and failing answers where criteria are subtle.
- Judge one criterion per call when accuracy matters. Multi-criterion prompts are cheaper but can blur criteria together.
Known biases and mitigations
| Bias | What happens | Mitigation |
|---|---|---|
| Position | In pairwise mode, prefers the first (or second) option | Run both orders, and count only consistent verdicts |
| Verbosity | Prefers longer answers | Instruct against it, compare at similar lengths, penalise padding in the rubric |
| Self-preference | Rates its own model family's outputs higher | Use a judge from a different family, or several judges |
| Leniency | Passes plausible-sounding wrong answers | Provide references or sources, and ask for specific evidence |
| Drift | Judge model updates change scores | Pin the judge model version and re-calibrate on upgrade |
Calibrating against humans
A judge is a measuring instrument, and you need to know its error:
- Have experts label a sample of 100–300 outputs with the same rubric.
- Run the judge on the same outputs.
- Measure agreement (accuracy, Cohen's kappa) and look at the disagreements. Fix the rubric or examples, and repeat.
- Only rely on the judge for gating once agreement is close to human-to-human agreement on the task.
- Periodically re-audit a sample, especially after changing judge model or rubric.
Online judging
The same judges can score live traffic samples for groundedness, policy compliance or user-intent success, feeding dashboards and alerts. Keep the judge off the request's critical path (run it asynchronously), unless it is acting as a guardrail, where latency budgets apply.
Key takeaways
- An LLM judge scores outputs against a rubric or reference. It scales human-like evaluation to thousands of cases and to live traffic.
- Specific, checkable criteria with a small scale or binary pass/fail beat vague 1–10 ratings.
- Judges have biases, including position, verbosity and self-preference. Mitigate them with swapped orders, length control and a different judge model.
- Calibrate against human labels, and trust a judge only when it agrees with experts at a known rate.