GenAI System Design
7. Evaluation, Observability and Guardrails

Why Evals Are the Test Suite

Why LLM systems need evaluation instead of (and as well as) unit tests, the kinds of evals, where they fit in the development loop, and how to explain an eval strategy in an interview.

Lesson 1 of 6 8 min

Why unit tests aren't enough

In a classic system, add(2, 3) returns 5 and a test asserts it. In an LLM system:

  • The same input can produce different, equally valid outputs.
  • "Correct" is often a judgement: helpful, faithful to the sources, the right tone.
  • A prompt change that fixes one case can quietly break ten others.
  • Model upgrades change behaviour across the board.

So instead of asserting exact outputs, you score outputs over a representative set of cases and track the scores over time. That set, plus its scoring, is an eval. Evals play the role for LLM systems that tests play for code, and teams that skip them end up changing prompts without knowing what broke.

Three levels of evaluation

Component evals
retrieval recall@k, classifier accuracy, schema validity
End-to-end evals
answer correctness, faithfulness, task success
Online evals
thumbs, edits, escalations, A/B tests
LevelMeasuresExample metricRuns
ComponentOne stage in isolationRecall@10 of retrieval, tool-selection accuracyEvery change to that stage
End-to-endThe whole pipeline's output% answers judged correct and groundedEvery prompt, model or pipeline change
OnlineReal users, real trafficThumbs-up rate, edit distance, resolution rate, conversionContinuously, and in A/B tests

Component evals tell you where a regression is. End-to-end evals tell you whether it matters. Online metrics tell you whether offline scores match reality.

What to measure

Choose metrics tied to the product's definition of good, from step 1 of the framework:

  • Correctness: matches a reference answer or the key facts.
  • Faithfulness / groundedness: every claim is supported by the retrieved sources.
  • Relevance / helpfulness: actually addresses the question.
  • Format compliance: valid schema, length, citations present.
  • Safety: no disallowed content, no PII leaks, resists injection attempts.
  • Task success for agents: the ticket was resolved, the tests pass, the right record changed.
  • Cost and latency per case, since quality alone can justify an expensive model.

Evals in the development loop

  1. Write eval cases before or alongside the feature, the way you would write tests.
  2. Run offline evals in CI on every prompt, model, retrieval or tool change. Block merges on regressions beyond a threshold.
  3. Compare variants on the same cases and inspect the cases that flipped, not just the average.
  4. Ship behind a flag or A/B test, and watch online metrics.
  5. Feed production failures back into the eval set, as new cases from thumbs-down, escalations and incident reviews.

This loop is what makes it safe to upgrade models, cut costs (can the smaller model pass?), and iterate on prompts without guessing.

Key takeaways

  • LLM outputs are probabilistic and open-ended, so you measure quality statistically over a set of cases instead of asserting exact outputs.
  • Evals come at three levels: component (retrieval recall, classification accuracy), end-to-end (answer quality, task success), and online (user signals, A/B tests).
  • Run offline evals on every prompt, model or pipeline change, like CI. Monitor online metrics after release.
  • Without evals you can't safely change prompts, upgrade models, or cut costs.

Go deeper

Finished reading? Mark it done to track your progress.