GenAI System Design
7. Evaluation, Observability and Guardrails

LLM-as-Judge

Using a language model to grade outputs at scale, covering rubric design, pointwise vs pairwise judging, known biases, calibrating against human labels, and cost.

Lesson 3 of 6 9 min

Why use a model to grade

Many qualities can't be checked with code: Is this answer faithful to the sources? Does it address the question? Is the tone right? Human review is accurate but costs minutes and dollars per case. An LLM judge reads the input, the output (and optionally a reference or sources), and returns a score with a rationale, in seconds and for cents or less.

Judge setups

SetupQuestion it answersUse for
Pointwise with rubric"Does this answer meet criteria X, Y, Z?"Absolute quality tracking, CI gates
Reference-based"Does this answer contain the facts in the reference?"Factual Q&A with known answers
Pairwise"Which of A and B is better?"Comparing prompt or model variants. More sensitive than pointwise
Groundedness / faithfulness"Is every claim supported by these sources?"RAG hallucination detection

Writing a good rubric

Vague: "Rate the answer's quality from 1 to 10." The scores are noisy and cluster around 7–8.

Better: break quality into specific, observable criteria, scored on a small scale or pass/fail:

Given the QUESTION, the SOURCES and the ANSWER, evaluate: 1. Correctness: Does the answer state the refund window accurately per the sources? (pass/fail) 2. Groundedness: Is every factual claim supported by a cited source? List any unsupported claims. (pass/fail) 3. Completeness: Does it address all parts of the question? (0 = no, 1 = partially, 2 = fully) First explain your reasoning briefly, then output JSON: {"correct": bool, "grounded": bool, "complete": 0|1|2}
  • Ask for reasoning before the verdict. It improves consistency.
  • Use structured output so scores are machine-readable.
  • Include examples of passing and failing answers where criteria are subtle.
  • Judge one criterion per call when accuracy matters. Multi-criterion prompts are cheaper but can blur criteria together.

Known biases and mitigations

BiasWhat happensMitigation
PositionIn pairwise mode, prefers the first (or second) optionRun both orders, and count only consistent verdicts
VerbosityPrefers longer answersInstruct against it, compare at similar lengths, penalise padding in the rubric
Self-preferenceRates its own model family's outputs higherUse a judge from a different family, or several judges
LeniencyPasses plausible-sounding wrong answersProvide references or sources, and ask for specific evidence
DriftJudge model updates change scoresPin the judge model version and re-calibrate on upgrade

Calibrating against humans

A judge is a measuring instrument, and you need to know its error:

  1. Have experts label a sample of 100–300 outputs with the same rubric.
  2. Run the judge on the same outputs.
  3. Measure agreement (accuracy, Cohen's kappa) and look at the disagreements. Fix the rubric or examples, and repeat.
  4. Only rely on the judge for gating once agreement is close to human-to-human agreement on the task.
  5. Periodically re-audit a sample, especially after changing judge model or rubric.

Online judging

The same judges can score live traffic samples for groundedness, policy compliance or user-intent success, feeding dashboards and alerts. Keep the judge off the request's critical path (run it asynchronously), unless it is acting as a guardrail, where latency budgets apply.

Key takeaways

  • An LLM judge scores outputs against a rubric or reference. It scales human-like evaluation to thousands of cases and to live traffic.
  • Specific, checkable criteria with a small scale or binary pass/fail beat vague 1–10 ratings.
  • Judges have biases, including position, verbosity and self-preference. Mitigate them with swapped orders, length control and a different judge model.
  • Calibrate against human labels, and trust a judge only when it agrees with experts at a known rate.

Go deeper

Finished reading? Mark it done to track your progress.