Skip to content
AI & Automation

AI Agent Observability: How to Monitor and Debug Agentic Systems

How to monitor and debug AI agents: traces and spans for model and tool calls, token and cost tracking, latency, errors, evaluation scores, OpenTelemetry GenAI conventions, privacy and incident investigation.

Quick answer

AI agent observability means recording every run as a trace: a root span for the run and child spans for each model call, tool call, retrieval and approval, carrying inputs, outputs, tool arguments, token counts, cost, latency, errors and version information. Aggregate these into metrics and dashboards, attach evaluation scores and user feedback, alert on errors, cost spikes, loops and quality drops, and protect the data, because traces contain customer information. Use the OpenTelemetry GenAI conventions to keep telemetry portable.

Where This Fits

Observability feeds agent evaluation and cost optimization, and records the decisions made by guardrails. General ecommerce observability practices are in ecommerce observability.

Model-level quality, drift and cost monitoring is covered in AI model monitoring.

Observability for non-agent LLM applications such as assistants and RAG systems is covered in LLM observability and tracing.

Why Agents Need Different Monitoring

A traditional service fails loudly: errors, timeouts, crashes. Agents often fail quietly: a plausible answer from the wrong source, an extra refund, a loop that burns tokens before giving up. To find these failures you need to see the content of each step and judge quality, not just measure availability.

Anatomy of an Agent Trace

Span typeKey attributes
Agent run (root)Run ID, user or tenant, task type, agent and prompt versions, outcome
Model callProvider, model, input and output tokens, latency, finish reason
Tool callTool name, arguments, result or error, latency, policy decision
RetrievalQuery, sources returned, scores, filters applied
ApprovalReviewer, decision, edits, wait time
Tool spans are where most consequential failures show up.

OpenTelemetry GenAI Conventions

The OpenTelemetry semantic conventions for generative AI define standard operation names (such as chat, invoke_agent and execute_tool) and attributes for providers, models and token usage. They are still evolving, but adopting them keeps telemetry portable between tools and lets traces from model SDKs, frameworks and MCP servers line up. The latest MCP specification also documents trace context propagation through its metadata fields.

Metrics and Dashboards

  • Runs per task type, success and escalation rates
  • Tokens and cost per run, per task type and per customer
  • Latency per step and end to end (p50 and p95)
  • Tool error rates and policy denials
  • Runs hitting step or cost limits
  • Online evaluation scores and user feedback
  • Model and prompt version distribution during rollouts

Agents in production but no idea what they are doing?

ZSpace Labs instruments agents with end-to-end tracing, cost tracking and quality alerts so issues are found before customers report them.

Start a Project

Alerts That Matter

Alert on changes that affect customers or cost: rising escalations or errors, sudden token or cost spikes, latency beyond budget, loops hitting limits, surges in policy denials (possible abuse or a broken tool) and drops in evaluation scores after a release or a provider model update. Tie alerts to owners and runbooks.

Debugging and Incident Investigation

When something goes wrong, the trace should answer: what did the agent receive, what did it retrieve, which tools did it call with which arguments, what came back, what policies decided and what it output. Replay the run against a fixed version to reproduce it. After the fix, add the case to the evaluation set. For incidents involving customers, record who was affected and what was changed so it can be corrected.

Privacy and Security of Telemetry

Traces hold prompts, documents and customer data. Redact or hash sensitive fields where possible, restrict access by role, encrypt at rest, set retention limits and exclude secrets entirely. If you use a third-party observability service, check where data is stored and how it is processed, and cover it in your privacy documentation.

Tooling Options

You can send OpenTelemetry traces to your existing observability platform, use LLM-specific observability tools that add prompt views, datasets and evaluations, or use your model provider's tracing features. Many teams combine a general platform for infrastructure with an LLM tool for content and quality. Choose based on data residency, cost and how well traces connect to evaluation.

How to Instrument an Agent Step by Step

  • 1. Create a run ID and propagate it through every service and tool
  • 2. Instrument model calls with model, tokens, latency and finish reason
  • 3. Instrument tool calls with arguments, results, errors and policy decisions
  • 4. Record versions of prompts, tools, models and retrieval indexes
  • 5. Add redaction and retention rules
  • 6. Build dashboards for success, cost, latency and errors
  • 7. Add alerts with owners and runbooks
  • 8. Connect traces to evaluation and feedback

What to Log and What Not To

DataLog it?Notes
Model, version, tokens, latency, costAlwaysCore operational data
Tool name, arguments, result statusAlwaysRedact sensitive argument values
Prompts and model outputsUsuallyRedact personal data; restrict access; set retention
Retrieved document IDs and scoresAlwaysIDs rather than full text where possible
Full retrieved textSometimesUseful for debugging; high privacy cost
Secrets, tokens, credentialsNeverStrip before logging
User and tenant identifiersAlwaysNeeded for access control and audits

Example Trace

A simplified trace shows how one run breaks down. Even this level of detail answers most debugging questions: where time went, which tool failed and how much the run cost.

Example: simplified trace of one agent run (illustrative)
invoke_agent  support_agent v12        run=run_91c  total=6.8s  cost=$0.031
├─ chat        model=small-model       in=1,820 out=96   0.9s  -> tool_call get_order
├─ execute_tool get_order(ORD-104233)                    0.3s  ok
├─ chat        model=small-model       in=2,410 out=88   0.8s  -> tool_call list_payments
├─ execute_tool list_payments(ORD-104233)                0.4s  ok (2 payments)
├─ chat        model=large-model       in=2,950 out=210  2.6s  -> tool_call refund_payment
├─ policy      refund_payment amount=59.90                       ALLOW (under auto limit)
├─ execute_tool refund_payment(pi_b, 59.90)              1.1s  ok
└─ chat        model=small-model       in=3,300 out=140  0.7s  -> final reply

Worked Example

An illustrative scenario, not a client case: costs for a research agent double overnight with no deploy. Traces show runs now average 18 steps instead of 7, because a search tool started returning errors that the agent retries repeatedly. The team adds a loop detector, a retry cap on that tool and an alert on step-count spikes, and asks the tool's owner to fix the error.

Common Mistakes

  • Logging only final answers
  • No cost attribution per run or customer
  • Storing full prompts with personal data indefinitely
  • No version information in traces
  • Alerts on uptime only

Want to see exactly why an agent did what it did?

Talk to ZSpace Labs about AI observability and agent operations and telemetry and backend integration.

Start a Project

Conclusion

You cannot run agents responsibly without seeing inside them. Trace every step, track cost and quality, alert on what matters and protect the data. Related: evaluation, LLM cost optimization and guardrails.

FAQ

Common questions

The ability to see what an agent did and why: traces of every model call, tool call, retrieval and approval in a run, with tokens, cost, latency, errors and quality scores, so problems can be detected and debugged.

Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.

Keep exploring
AI & Automation
7 min read

AI Agent Evaluation: How to Test Accuracy, Reliability and Performance

How to evaluate AI agents: building evaluation datasets, task success, tool-call accuracy, groundedness, policy compliance, latency, cost, LLM-as-judge, regression testing and production evaluation.

Read article
AI & Automation
7 min read

LLM Cost Optimization: How to Control the Cost of AI Applications

How to reduce the cost of LLM applications without losing quality: measuring cost per task, trimming context, output limits, model routing, prompt and response caching, batch processing, agent step budgets and governance.

Read article
AI & Automation
7 min read

AI Agent Guardrails: How to Control What Autonomous Agents Can Do

How to put guardrails on AI agents: permission boundaries, tool restrictions, input and output validation, policy engines, action approvals, rate limits and safe execution for autonomous systems.

Read article