Agent observability
Agent observability is the practice of collecting and interpreting evidence about an AI agent's multi-step execution: model calls, decisions, tool invocations, retrieval, state changes, latency, cost, errors and outcomes. It connects events into a trajectory so an operator can reconstruct what the system attempted and where it failed. Logging a single prompt and response is therefore insufficient for an agent with tools and memory.
Origin and context
Application performance monitoring and distributed tracing provide the technical ancestry. Arize's 2023 OpenInference explanation described traces joining model calls, retrieval and tools. The November 2024 AgentOps preprint organized artifacts across an agent lifecycle; it is a research taxonomy, not a completed industry standard. A dated 2025 OpenTelemetry account distinguished application instrumentation from framework conventions. AgentOps is the broader lifecycle practice, tracing is a mechanism, and semantic conventions describe the exchanged data. None is an alias for observability itself.
Why it matters
Agent failures can emerge from interactions between components rather than one incorrect final answer. A retrieval step can supply irrelevant context that a later model faithfully summarizes. Linked observations let an engineer move from the unsuccessful result to the participating calls, or from an anomalous component to affected runs. They also support comparison of latency, token use and evaluation results across runs. Skills Intelligence treats this connection between execution evidence and quality assessment as the defining operational value, not the volume of logs collected.
Example
Consider a research assistant that searches a document collection and summarizes the results. A linked trace records the request, retrieved document identifiers, model call, duration and final output. If the summary is outdated, the engineer can inspect the retrieval span to check whether the agent received obsolete material. This is an illustrative diagnostic workflow, not proof that the trace identifies the sole cause: the retrieval query, indexing process and final response still need separate evaluation.
Maturity and evidence
Maturity is rated 3. An agent-specific research taxonomy, an independent OpenTelemetry initiative and Arize's implemented tracing approach establish a recognizable practice across organizations. The evidence supports the practice, not one universally adopted agent schema. The rating remains below 4 because these sources describe different instrumentation boundaries and evaluation methods; telemetry compatibility and behavioral quality remain separate questions.
Limits and open questions
A trace records the operations that were instrumented; it is not a complete account of a model's internal computation or a causal proof. An absent span may mean missing instrumentation rather than an absent action. Likewise, short latency and successful tool responses do not establish that the overall task was completed correctly. Observability therefore supplies evidence for evaluation and troubleshooting, rather than replacing either.
Related terms
References
- A Taxonomy of AgentOps for Enabling Observability of Foundation Model based Agents (v1 preprint)Dong, Lu and Zhu / CSIRO Data61 and UNSW · 2024-11-08 · class A
- AI Agent Observability - Evolving Standards and Best PracticesOpenTelemetry / Guangya Liu (IBM) and Sujay Solomon (Google) · 2025-03-06 · class A
- LLM Observability: One Small Step for Spans, One Giant Leap for Span-KindsArize AI / Amber Roberts · 2023-09-27 · class A
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as llm observability skill.
This term is also covered in the Skills Atlas as agent evaluation skill.
This term is also covered in the Skills Atlas as ml monitoring skill.