AI Observability
Instrument AI and LLM applications with OpenTelemetry to get unified traces that connect HTTP requests, agent orchestration, LLM API calls, and database queries in a single view.
The Problem
Traditional APM tools (Datadog, New Relic) capture HTTP and database telemetry. Specialized AI tools (LangSmith, Weights & Biases) capture LLM traces. Neither shows the full picture:
| Tool Type | Captures | Misses |
|---|---|---|
| Traditional APM | HTTP requests, DB queries, latency | Model name, tokens, cost, prompt content |
| AI-specific tools | LLM calls, prompts, model metadata | HTTP context, DB queries, infrastructure |
| OpenTelemetry | All of the above in one trace | - |
With OpenTelemetry, a single trace shows that a slow HTTP response was caused by a specific LLM call in a specific agent, which also triggered 3 database queries and a fallback to a different provider. Keeping APM and LLM telemetry in one backend also avoids paying for two stacks; see LLM observability cost for the numbers.
When to Use AI Observability
| Use Case | Recommendation |
|---|---|
| Track LLM token usage and costs | AI Observability |
| Monitor agent pipeline performance | AI Observability |
| Evaluate LLM output quality over time | AI Observability |
| Debug slow AI requests end-to-end | AI Observability |
| Attribute costs to agents or business operations | AI Observability |
| Standard HTTP/database monitoring only | Auto-instrumentation |
| Generic custom spans and metrics | Custom instrumentation |
Guides
| Guide | What It Covers |
|---|---|
| AI Agent Observability | Framework-agnostic concepts and patterns: the agent timeline, GenAI operation types, conversation-id propagation, tool-call (execute_tool) instrumentation, MCP tools, multi-agent handoffs, agent metrics and evaluation |
| Agent Approval Gates | Human-in-the-loop agents (C#, Microsoft Agent Framework): approval as a linked span pair plus a wait histogram, MCP tool-call spans and params._meta trace context, trace context across the pause, and which span carries the error status |
| LLM Observability | End-to-end guide (Python): GenAI semantic conventions, token/cost metrics, agent pipeline spans, evaluation tracking, PII scrubbing, production deployment |
| Rust LLM Observability | End-to-end guide (Rust): GenAI semantic conventions, multi-provider LLM with fallback, token/cost metrics, multi-stage pipeline spans, retry observability, Docker deployment |
| Spring AI LLM Observability | End-to-end guide (Java): Three-layer instrumentation (Java Agent + Spring AI + manual OTel), GenAI semantic conventions, tool calling, RAG, domain metrics, Docker deployment |
| LangChain Instrumentation | Framework-specific: the official OpenTelemetry GenAI instrumentation for LangChain, agent, chat, tool and retrieval spans, conversation IDs, cost and scrubbing in a span exporter, retries in middleware and known gaps |
| LangChain Callback Handler | Writing your own LangChain callback handler: run tree to span tree, parenting on parent_run_id, error handling, for chains the official package does not cover |
| LangGraph Instrumentation | Framework-specific: LangGraph node wrapping, conditional edge routing, tool-calling nodes, state management, pipeline traces |
| LlamaIndex Instrumentation | Framework-specific: LlamaIndex model calls traced by the official OpenTelemetry GenAI SDK packages, Ollama through OpenAILike, request context, cost and scrubbing in a span processor and exporter, structured output corrections and known gaps |
| Vercel AI SDK Instrumentation | Framework-specific: AI SDK 7 agent spans via @ai-sdk/otel, run ids, per-run cost and subagent fan-out, plus the v6 middleware path |
| Pydantic AI Instrumentation | Framework-specific: Pydantic AI's built-in OpenTelemetry via Agent.instrument_all, agent, model and tool spans, token metrics, content capture and prompt versions, no Logfire |
| Pydantic AI on Temporal | Framework-specific: one trace per durable Pydantic AI workflow on Temporal, across replay and worker restarts, with replay-safe logs and metrics |
| Strands Agents Instrumentation | Framework-specific: Strands Agents' built-in OpenTelemetry, agent, model and tool spans, an agent called as a tool, trace-correlated logs, redaction and known gaps |
| Google ADK Instrumentation | Framework-specific: Google ADK's built-in OpenTelemetry, agent, model and tool spans, gen_ai.* metrics and inference events, an agent called through AgentTool, content capture and known gaps |
| Microsoft Agent Framework Instrumentation | Framework-specific: Agent Framework's built-in OpenTelemetry for Python and .NET, agent, chat and tool spans, gen_ai.* metrics and message events, sensitive data, middleware budgets and known gaps |
| OpenAI Agents SDK Instrumentation | Framework-specific: the contrib OpenAI Agents and OpenAI instrumentations, workflow, agent, chat and tool spans, gen_ai.* metrics, trace export off OpenAI, content capture modes and known gaps |
| Mastra Instrumentation | Framework-specific: Mastra's tracing on the OpenTelemetry SDK through @mastra/otel-bridge, agent, chat and tool spans in the request's trace, an agent called from a tool, content hiding, your IDs in a span processor and known gaps |
| OpenClaw | Agent runtime: the gateway's diagnostics-otel plugin, run, model and tool spans, gen_ai.* and openclaw.* metrics, trace-correlated logs and known gaps |
| Claude Code | Agent runtime: Claude Code and the Claude Agent SDK by environment variables, interaction, model request and tool spans, cost and token metrics, per-event logs and known gaps |
| Codex CLI | Agent runtime: the OpenAI Codex CLI's [otel] config, turn, model request and command spans, codex.* metrics, per-event logs, a stream event filter and known gaps |
What Gets Instrumented
AI observability builds on top of auto and custom instrumentation, adding an LLM-specific layer:
Auto-Instrumentation Layer (zero code changes)
- HTTP requests via FastAPI/Django/Flask instrumentors (Python), tower-http TraceLayer (Rust), Java Agent (Spring WebFlux)
- Database queries via SQLAlchemy/Django ORM instrumentors (Python), SQLx tracing (Rust), Java Agent (JDBC/R2DBC)
- Outbound HTTP via httpx/requests instrumentors (Python), Java Agent (captures raw LLM API calls)
- Log correlation via logging instrumentor (Python), OpenTelemetryTracingBridge (Rust), Java Agent (Logback/Log4j)
Custom AI Layer (GenAI semantic conventions)
- LLM spans with model, provider, token counts, cost
- Prompt/completion events with PII scrubbing
- Agent spans with pipeline orchestration context
- Tool-call spans (
execute_tool) with tool name, arguments, result, and error type - where most agentic failures live - Multi-agent handoffs bound by
gen_ai.conversation.idso a full run reads as one timeline across agents and downstream services - Evaluation events with quality scores and pass/fail
- Cost metrics with attribution by agent and business operation
- Retry/fallback tracking with error type classification
Example: Unified Trace
POST /api/generate 4.2s [auto: HTTP]
├─ db.query SELECT context 15ms [auto: DB]
├─ invoke_agent enrich 1.8s [custom: agent]
│ └─ gen_ai.chat claude-sonnet-4 1.7s [custom: LLM]
│ └─ HTTP POST api.anthropic.com 1.7s [auto: httpx]
├─ invoke_agent draft 2.3s [custom: agent]
│ └─ gen_ai.chat claude-sonnet-4 2.2s [custom: LLM]
│ └─ HTTP POST api.anthropic.com 2.2s [auto: httpx]
└─ db.query INSERT result 5ms [auto: DB]
Key Concepts
GenAI Semantic Conventions
OpenTelemetry defines GenAI semantic conventions for standardized LLM telemetry. Key attributes:
| Attribute | Example | Purpose |
|---|---|---|
gen_ai.operation.name | "invoke_agent" | Operation type |
gen_ai.provider.name | "anthropic" | LLM provider |
gen_ai.request.model | "claude-sonnet-4" | Model used |
gen_ai.usage.input_tokens | 1240 | Tokens consumed |
gen_ai.usage.output_tokens | 320 | Tokens generated |
gen_ai.conversation.id | "conv-8f2a" | Binds a full agent run |
gen_ai.agent.name | "draft" | Agent in pipeline |
gen_ai.tool.name | "search_orders" | Tool an agent invoked |
For the agent-specific attributes (gen_ai.conversation.id,
gen_ai.tool.*, multi-agent handoffs), see the
AI Agent Observability guide.
GenAI Metrics
Custom metrics for dashboards and alerting:
| Metric | Type | Purpose |
|---|---|---|
gen_ai.client.token.usage | Histogram | Token consumption by model/agent |
gen_ai.client.operation.duration | Histogram | LLM call latency |
gen_ai.client.cost | Counter | Cost in USD by model/agent |
gen_ai.evaluation.score | Histogram | Output quality scores |
gen_ai.client.error.count | Counter | Errors by provider/type |
Next Steps
- Follow the LLM Observability guide for a complete Python setup walkthrough
- Follow the Rust LLM Observability guide for Rust AI applications with manual GenAI instrumentation
- Follow the Spring AI LLM Observability guide for Java Spring AI applications with three-layer instrumentation
- Set up auto-instrumentation for your web framework if you haven't already
- Configure the OpenTelemetry Collector to export telemetry to base14 Scout