GenAI System Design
7. Evaluation, Observability and Guardrails

Tracing and Observability

What to log for every LLM request, how to trace multi-step RAG and agent pipelines, the dashboards and alerts that matter, and handling the privacy cost of prompt logging.

Lesson 4 of 6 9 min

Why LLM observability is different

Classic observability asks "is it up and fast?" LLM systems add:

  • Is it good? An HTTP 200 response can be a confident hallucination.
  • What did it cost? Cost per request varies 100×.
  • Why did it do that? The answer depends on the exact prompt, retrieved chunks and tool results, which you can't reconstruct afterwards unless you logged them.

Traces for multi-step pipelines

Use distributed tracing (OpenTelemetry has GenAI semantic conventions) with one span per step:

Request span
user, tenant, feature
Retrieval span
query, top-k IDs + scores
LLM span
model, tokens, TTFT, cost
Tool spans
name, args, duration, result size
Guardrail spans
verdicts
Agents produce trees: nested LLM and tool spans per step, and child traces per sub-agent.

With this, you can open any bad answer and see exactly which chunks were retrieved, what prompt was sent, how long each step took, and what it cost.

What to log per LLM call

FieldWhy
Trace and span ID, parent IDReconstruct pipelines and agent trees
Feature, tenant, user (pseudonymous)Attribution, per-tenant debugging
Model + version, provider, regionDetect provider-side changes, compare models
Prompt template ID + versionTie behaviour to prompt changes
Input, cached and output tokensCost, cache hit rate, context growth
TTFT, total latency, TPOTLatency SLOs
Cost (computed)Budgets, chargeback
Finish reason (stop, length, tool, filter)Truncation and refusal detection
Errors, retries, fallback usedReliability
Prompt and response content (redacted, sampled)Debugging, building eval cases

Dashboards and alerts

Quality

  • Online judge scores on sampled traffic: groundedness, policy compliance
  • User feedback rate and thumbs-down reasons, edit rates, escalations
  • "No answer found" rate (content gaps), refusal rate

Latency

  • TTFT p50/p95 and total p95, per feature and model, bucketed by prompt size

Cost

  • Spend per day by feature, tenant and model, and cost per task p50/p95 for agents
  • Cache hit rate (a drop usually means someone broke the prompt prefix)

Reliability

  • Error rate, 429 rate, fallback rate, circuit breaker state
  • Finish reason length rate (outputs being truncated)

Alert on changes: a sudden rise in cost per request, a drop in cache hit rate, a spike in refusals, or a TTFT p95 breach. Many LLM incidents are silent quality or cost regressions, not outages.

Closing the loop

Observability data feeds the eval process:

  • Thumbs-down traces go to a review queue, then to labelled eval cases.
  • Clusters of similar failing queries point to a retrieval gap or a missing tool.
  • Cost outliers point to runaway agent loops or prompt bloat.

Privacy and retention

Prompts and outputs contain user data, often sensitive:

  • Redact PII before storage, or store content in a separate, more restricted system than metrics.
  • Sample content logging (for example 5%), and keep metrics at 100%.
  • Restrict access to raw traces, and audit who views them.
  • Set retention limits (for example 30 days for content, longer for metrics), and honour deletion requests across trace stores too.
  • For multi-tenant systems, scope trace access by tenant.

Key takeaways

  • Trace every request end to end, with spans for retrieval, each LLM call, each tool call and guardrails, linked by one trace ID.
  • For each LLM span, log model, prompt version, token counts (input, cached, output), TTFT, total latency, cost, finish reason and errors.
  • Dashboards cover quality (online scores, feedback), latency (TTFT and total p95), cost (per feature, tenant and task) and reliability (errors, 429s, fallbacks).
  • Prompts and outputs contain user data. Redact, restrict access, and set retention limits.

Go deeper

Finished reading? Mark it done to track your progress.