Tracing and Observability
What to log for every LLM request, how to trace multi-step RAG and agent pipelines, the dashboards and alerts that matter, and handling the privacy cost of prompt logging.
Why LLM observability is different
Classic observability asks "is it up and fast?" LLM systems add:
- Is it good? An HTTP 200 response can be a confident hallucination.
- What did it cost? Cost per request varies 100×.
- Why did it do that? The answer depends on the exact prompt, retrieved chunks and tool results, which you can't reconstruct afterwards unless you logged them.
Traces for multi-step pipelines
Use distributed tracing (OpenTelemetry has GenAI semantic conventions) with one span per step:
With this, you can open any bad answer and see exactly which chunks were retrieved, what prompt was sent, how long each step took, and what it cost.
What to log per LLM call
| Field | Why |
|---|---|
| Trace and span ID, parent ID | Reconstruct pipelines and agent trees |
| Feature, tenant, user (pseudonymous) | Attribution, per-tenant debugging |
| Model + version, provider, region | Detect provider-side changes, compare models |
| Prompt template ID + version | Tie behaviour to prompt changes |
| Input, cached and output tokens | Cost, cache hit rate, context growth |
| TTFT, total latency, TPOT | Latency SLOs |
| Cost (computed) | Budgets, chargeback |
| Finish reason (stop, length, tool, filter) | Truncation and refusal detection |
| Errors, retries, fallback used | Reliability |
| Prompt and response content (redacted, sampled) | Debugging, building eval cases |
Dashboards and alerts
Quality
- Online judge scores on sampled traffic: groundedness, policy compliance
- User feedback rate and thumbs-down reasons, edit rates, escalations
- "No answer found" rate (content gaps), refusal rate
Latency
- TTFT p50/p95 and total p95, per feature and model, bucketed by prompt size
Cost
- Spend per day by feature, tenant and model, and cost per task p50/p95 for agents
- Cache hit rate (a drop usually means someone broke the prompt prefix)
Reliability
- Error rate, 429 rate, fallback rate, circuit breaker state
- Finish reason
lengthrate (outputs being truncated)
Alert on changes: a sudden rise in cost per request, a drop in cache hit rate, a spike in refusals, or a TTFT p95 breach. Many LLM incidents are silent quality or cost regressions, not outages.
Closing the loop
Observability data feeds the eval process:
- Thumbs-down traces go to a review queue, then to labelled eval cases.
- Clusters of similar failing queries point to a retrieval gap or a missing tool.
- Cost outliers point to runaway agent loops or prompt bloat.
Privacy and retention
Prompts and outputs contain user data, often sensitive:
- Redact PII before storage, or store content in a separate, more restricted system than metrics.
- Sample content logging (for example 5%), and keep metrics at 100%.
- Restrict access to raw traces, and audit who views them.
- Set retention limits (for example 30 days for content, longer for metrics), and honour deletion requests across trace stores too.
- For multi-tenant systems, scope trace access by tenant.
Key takeaways
- Trace every request end to end, with spans for retrieval, each LLM call, each tool call and guardrails, linked by one trace ID.
- For each LLM span, log model, prompt version, token counts (input, cached, output), TTFT, total latency, cost, finish reason and errors.
- Dashboards cover quality (online scores, feedback), latency (TTFT and total p95), cost (per feature, tenant and task) and reliability (errors, 429s, fallbacks).
- Prompts and outputs contain user data. Redact, restrict access, and set retention limits.