LLM Observability: How to Monitor AI Application Quality and Performance
How to observe LLM applications in production: traces and spans across retrieval, model and tool calls, correlation IDs, token usage, cost, latency, errors, retrieval quality, output quality, user feedback and debugging multi-step workflows.
Quick answer
LLM observability means capturing enough data to explain any production response: a trace for every request with spans for retrieval, prompt assembly, model calls, tool calls and validation, each carrying inputs, outputs, versions, tokens, latency and errors, joined by correlation IDs. On top of traces, track metrics for latency, error rates, cost and cache use, plus quality signals from sampled evaluation and user feedback. Protect the data, because traces often contain sensitive content.
Where This Fits
This article covers observability and tracing for LLM applications such as assistants, RAG systems and AI features. Agent-specific concerns are in AI agent observability, model quality and drift in AI model monitoring, and the wider operating practice in LLMOps.
Why Traditional Monitoring Falls Short
A conventional dashboard might show a support assistant with a 99.9% success rate and healthy latency while it confidently gives customers an outdated refund policy. Nothing failed technically: retrieval returned an old document, the model used it faithfully, and the response was well formed. Only by seeing the retrieved context, prompt version and output together can you find the cause.
LLM applications also have costs that vary per request, behaviour that changes when a provider updates a model, and multi-step flows where one weak step spoils the result. Observability for these systems therefore combines classic telemetry with content, versions and quality.
Anatomy of an LLM Trace
A trace represents one user request. Inside it, spans represent steps, nested where one step calls another. Every span carries timing, status and attributes; LLM spans add model, prompt version, token counts and, where permitted, input and output content.
trace_id: 7f3c... user: u_482 (hashed) feature: support_answer release: 2026.10.2
├─ span api.request 1,842 ms status=ok
│ ├─ span retrieval.search 146 ms index=kb_v14 top_k=8 hits=8
│ ├─ span prompt.build 3 ms prompt=support_answer@v12
│ ├─ span llm.chat 1,512 ms model=<provider/model> in=3,904 tok out=412 tok
│ │ cost=$0.0071 finish=stop
│ ├─ span output.validate 11 ms schema=ok citations=3/3 found
│ └─ span response.stream 165 ms
feedback: thumbs_down reason="policy outdated"Standardizing With OpenTelemetry
The OpenTelemetry semantic conventions for generative AI define standard attribute names for model calls, such as the provider, requested and response model, token usage and operation type. Instrumenting against these conventions lets you send the same data to different backends and combine AI spans with the rest of your distributed tracing.
The conventions are still evolving, so pin library versions and expect some attribute changes over time. Many frameworks and SDKs offer automatic instrumentation, but check what content they capture by default before enabling them in production.
What to Measure
| Category | Metrics | Typical alert |
|---|---|---|
| Latency | Time to first token, total time, p50 and p95 per feature | p95 above budget |
| Errors | Provider errors, timeouts, rate limits, validation failures | Error rate above baseline |
| Usage and cost | Input and output tokens, cost per request and per feature, cache hit rate | Daily spend or cost per request spike |
| Retrieval | Hit counts, empty results, score distributions, source freshness | Rising empty-result rate |
| Quality | Sampled judge scores, feedback ratio, edits, retries, escalations | Score drop after a release |
| Safety | Policy flags, injection detections, blocked tool calls | Any spike or new pattern |
Can't tell why your AI feature gives bad answers?
ZSpace Labs instruments LLM applications end to end so every answer can be explained. See our AI engineering services.
Debugging Multi-Step Workflows With Traces
Good traces turn vague complaints into specific causes. A practical routine: find the trace from the user's report or feedback, check retrieval first (were the right documents found, and were they current?), then the assembled prompt (was the context included and the right prompt version used?), then the model output (did it ignore or misread the context?), then validation and post-processing.
For workflows that span services, queues or external APIs, propagate the trace context through every hop, including message headers on queues, so asynchronous steps join the same trace. Where a hop cannot carry trace context, log a correlation ID that links the records. Patterns that recur, such as a document source that is often stale, become fixes and new evaluation cases.
Quality Signals Without Labels
Most production outputs never receive a ground-truth label, so combine several signals. Explicit feedback is valuable but sparse and skewed toward strong reactions. Implicit signals include users editing drafts heavily, retrying, abandoning or escalating to a person. Sampled evaluation scores a small fraction of traces with the same judges used before release, giving a consistent trend. Watch all three by release version so you can tell whether a change helped. Feedback design is covered in AI feedback UX.
Protecting Trace Data
Traces can become the largest store of sensitive data in an AI system: user questions, retrieved documents and model outputs. Redact or hash identifiers, avoid capturing secrets and payment data, restrict trace access by role, set short retention for full content and keep aggregate metrics longer. Check what your instrumentation captures by default. Privacy guidance is in AI data privacy and leakage risks in AI data leakage.
Choosing Tooling
Teams typically choose between dedicated LLM observability platforms, which provide trace views of prompts and outputs, evaluation and feedback features, and general observability backends that receive OpenTelemetry data alongside the rest of the system. Examples of LLM-focused tools include Langfuse, LangSmith and MLflow tracing. Consider data residency and self-hosting options, cost at your trace volume and how well the tool fits your evaluation workflow.
Advantages and Limitations
Observability shortens incident investigations from days to minutes, makes cost visible per feature and shows whether releases improve quality. Its limits: storing content is expensive and sensitive, sampled quality scores are estimates and instrumentation adds some overhead and maintenance as conventions change.
How to Set Up Observability Step by Step
- 1. Instrument every model, retrieval and tool call as spans with OpenTelemetry
- 2. Attach versions (release, prompt, model, index) to every trace
- 3. Propagate trace context across services and queues
- 4. Record tokens and cost and build per-feature dashboards
- 5. Add feedback capture linked to trace IDs
- 6. Score a sample of traces with your evaluation judges
- 7. Set alerts for errors, latency, cost and quality drops
- 8. Apply redaction, access control and retention to trace data
Sampling and Retention Strategy
Capturing every request in full detail is expensive and multiplies privacy risk. A common approach: keep metrics and span metadata (timings, token counts, versions, status) for all requests; keep full content for a sample, for all requests with negative feedback or errors, and for flagged policy events; and keep full content only for a short period unless needed for an investigation or evaluation set.
Tail-based sampling, where the decision to keep a trace is made after it completes, lets you keep all slow, failed or flagged traces while sampling normal ones. Document what is captured and for how long, and align it with your privacy notices.
Observability for Cost Governance
Because every model span records tokens and cost, traces become the most accurate source for AI spend by feature, team, customer or tenant. Tag requests with feature and tenant identifiers, build dashboards for cost per request and per successful task, and alert on sudden changes, such as a prompt edit that doubled context size. These views turn cost discussions from guesses into data. Cost levers are covered in LLM cost optimization, and gateway-level attribution in AI platform engineering.
Example Instrumentation Attributes
Whichever tools you use, agree a small, consistent set of attributes on every AI span so dashboards and queries work across features. Align names with the OpenTelemetry generative AI conventions where they exist, and add your own for product context.
| Attribute | Example | Purpose |
|---|---|---|
| Provider and model | provider, requested and response model | Compare versions, detect silent changes |
| Token usage | input and output tokens, cached tokens | Cost and efficiency |
| Prompt version | support_answer@v12 | Tie behaviour to releases |
| Feature and tenant | feature=support_answer, tenant=t_93 | Attribution and isolation checks |
| Retrieval details | index version, chunk IDs, scores | Debug grounding |
| Outcome | validation status, finish reason, feedback | Quality signals |
Worked Example
An illustrative scenario, not a client case: an HR assistant's thumbs-down rate doubles in a week with no errors logged. Filtering traces with negative feedback shows most involve leave questions, and the retrieval spans show an archived 2024 leave policy ranking above the current one after a re-index. The team excludes archived documents from the index, adds the failing questions to the evaluation set and adds an alert on retrieval of documents past their review date.
Common Mistakes
- Logging only errors and latency, not content and versions
- Traces that stop at service boundaries or queues
- Capturing full prompts with personal data and keeping them indefinitely
- No link between user feedback and the trace it refers to
- Dashboards nobody owns or reviews
Want observability that explains every AI answer?
Talk to ZSpace Labs about LLM tracing and monitoring built on open standards and your existing stack.
Conclusion
LLM observability gives you the evidence to answer why a response happened, what it cost and whether quality is improving. Trace every step with versions, measure cost and quality alongside latency, protect the data and turn what you find into tests.
Common questions
The ability to understand what an LLM application did and why, from data it emits in production: traces of each request's steps, metrics for latency, errors, tokens and cost, logs, and quality signals such as evaluation scores and user feedback.