Skip to main content

AI Observability

Instrument AI and LLM applications with OpenTelemetry to get unified traces that connect HTTP requests, agent orchestration, LLM API calls, and database queries in a single view.

The Problem​

Traditional APM tools (Datadog, New Relic) capture HTTP and database telemetry. Specialized AI tools (LangSmith, Weights & Biases) capture LLM traces. Neither shows the full picture:

Tool TypeCapturesMisses
Traditional APMHTTP requests, DB queries, latencyModel name, tokens, cost, prompt content
AI-specific toolsLLM calls, prompts, model metadataHTTP context, DB queries, infrastructure
OpenTelemetryAll of the above in one trace-

With OpenTelemetry, a single trace shows that a slow HTTP response was caused by a specific LLM call in a specific agent, which also triggered 3 database queries and a fallback to a different provider. Keeping APM and LLM telemetry in one backend also avoids paying for two stacks; see LLM observability cost for the numbers.

When to Use AI Observability​

Use CaseRecommendation
Track LLM token usage and costsAI Observability
Monitor agent pipeline performanceAI Observability
Evaluate LLM output quality over timeAI Observability
Debug slow AI requests end-to-endAI Observability
Attribute costs to agents or business operationsAI Observability
Standard HTTP/database monitoring onlyAuto-instrumentation
Generic custom spans and metricsCustom instrumentation

Guides​

GuideWhat It Covers
AI Agent ObservabilityFramework-agnostic concepts and patterns: the agent timeline, GenAI operation types, conversation-id propagation, tool-call (execute_tool) instrumentation, MCP tools, multi-agent handoffs, agent metrics and evaluation
Agent Approval GatesHuman-in-the-loop agents (C#, Microsoft Agent Framework): approval as a linked span pair plus a wait histogram, MCP tool-call spans and params._meta trace context, trace context across the pause, and which span carries the error status
LLM ObservabilityEnd-to-end guide (Python): GenAI semantic conventions, token/cost metrics, agent pipeline spans, evaluation tracking, PII scrubbing, production deployment
Rust LLM ObservabilityEnd-to-end guide (Rust): GenAI semantic conventions, multi-provider LLM with fallback, token/cost metrics, multi-stage pipeline spans, retry observability, Docker deployment
Spring AI LLM ObservabilityEnd-to-end guide (Java): Three-layer instrumentation (Java Agent + Spring AI + manual OTel), GenAI semantic conventions, tool calling, RAG, domain metrics, Docker deployment
LangChain InstrumentationFramework-specific: the official OpenTelemetry GenAI instrumentation for LangChain, agent, chat, tool and retrieval spans, conversation IDs, cost and scrubbing in a span exporter, retries in middleware and known gaps
LangChain Callback HandlerWriting your own LangChain callback handler: run tree to span tree, parenting on parent_run_id, error handling, for chains the official package does not cover
LangGraph InstrumentationFramework-specific: LangGraph node wrapping, conditional edge routing, tool-calling nodes, state management, pipeline traces
LlamaIndex InstrumentationFramework-specific: LlamaIndex model calls traced by the official OpenTelemetry GenAI SDK packages, Ollama through OpenAILike, request context, cost and scrubbing in a span processor and exporter, structured output corrections and known gaps
Vercel AI SDK InstrumentationFramework-specific: AI SDK 7 agent spans via @ai-sdk/otel, run ids, per-run cost and subagent fan-out, plus the v6 middleware path
Pydantic AI InstrumentationFramework-specific: Pydantic AI's built-in OpenTelemetry via Agent.instrument_all, agent, model and tool spans, token metrics, content capture and prompt versions, no Logfire
Pydantic AI on TemporalFramework-specific: one trace per durable Pydantic AI workflow on Temporal, across replay and worker restarts, with replay-safe logs and metrics
Strands Agents InstrumentationFramework-specific: Strands Agents' built-in OpenTelemetry, agent, model and tool spans, an agent called as a tool, trace-correlated logs, redaction and known gaps
Google ADK InstrumentationFramework-specific: Google ADK's built-in OpenTelemetry, agent, model and tool spans, gen_ai.* metrics and inference events, an agent called through AgentTool, content capture and known gaps
Microsoft Agent Framework InstrumentationFramework-specific: Agent Framework's built-in OpenTelemetry for Python and .NET, agent, chat and tool spans, gen_ai.* metrics and message events, sensitive data, middleware budgets and known gaps
OpenAI Agents SDK InstrumentationFramework-specific: the contrib OpenAI Agents and OpenAI instrumentations, workflow, agent, chat and tool spans, gen_ai.* metrics, trace export off OpenAI, content capture modes and known gaps
Mastra InstrumentationFramework-specific: Mastra's tracing on the OpenTelemetry SDK through @mastra/otel-bridge, agent, chat and tool spans in the request's trace, an agent called from a tool, content hiding, your IDs in a span processor and known gaps
OpenClawAgent runtime: the gateway's diagnostics-otel plugin, run, model and tool spans, gen_ai.* and openclaw.* metrics, trace-correlated logs and known gaps
Claude CodeAgent runtime: Claude Code and the Claude Agent SDK by environment variables, interaction, model request and tool spans, cost and token metrics, per-event logs and known gaps
Codex CLIAgent runtime: the OpenAI Codex CLI's [otel] config, turn, model request and command spans, codex.* metrics, per-event logs, a stream event filter and known gaps

What Gets Instrumented​

AI observability builds on top of auto and custom instrumentation, adding an LLM-specific layer:

Auto-Instrumentation Layer (zero code changes)​

  • HTTP requests via FastAPI/Django/Flask instrumentors (Python), tower-http TraceLayer (Rust), Java Agent (Spring WebFlux)
  • Database queries via SQLAlchemy/Django ORM instrumentors (Python), SQLx tracing (Rust), Java Agent (JDBC/R2DBC)
  • Outbound HTTP via httpx/requests instrumentors (Python), Java Agent (captures raw LLM API calls)
  • Log correlation via logging instrumentor (Python), OpenTelemetryTracingBridge (Rust), Java Agent (Logback/Log4j)

Custom AI Layer (GenAI semantic conventions)​

  • LLM spans with model, provider, token counts, cost
  • Prompt/completion events with PII scrubbing
  • Agent spans with pipeline orchestration context
  • Tool-call spans (execute_tool) with tool name, arguments, result, and error type - where most agentic failures live
  • Multi-agent handoffs bound by gen_ai.conversation.id so a full run reads as one timeline across agents and downstream services
  • Evaluation events with quality scores and pass/fail
  • Cost metrics with attribution by agent and business operation
  • Retry/fallback tracking with error type classification

Example: Unified Trace​

Single trace spanning all layers
POST /api/generate 4.2s [auto: HTTP]
├─ db.query SELECT context 15ms [auto: DB]
├─ invoke_agent enrich 1.8s [custom: agent]
│ └─ gen_ai.chat claude-sonnet-4 1.7s [custom: LLM]
│ └─ HTTP POST api.anthropic.com 1.7s [auto: httpx]
├─ invoke_agent draft 2.3s [custom: agent]
│ └─ gen_ai.chat claude-sonnet-4 2.2s [custom: LLM]
│ └─ HTTP POST api.anthropic.com 2.2s [auto: httpx]
└─ db.query INSERT result 5ms [auto: DB]

Key Concepts​

GenAI Semantic Conventions​

OpenTelemetry defines GenAI semantic conventions for standardized LLM telemetry. Key attributes:

AttributeExamplePurpose
gen_ai.operation.name"invoke_agent"Operation type
gen_ai.provider.name"anthropic"LLM provider
gen_ai.request.model"claude-sonnet-4"Model used
gen_ai.usage.input_tokens1240Tokens consumed
gen_ai.usage.output_tokens320Tokens generated
gen_ai.conversation.id"conv-8f2a"Binds a full agent run
gen_ai.agent.name"draft"Agent in pipeline
gen_ai.tool.name"search_orders"Tool an agent invoked

For the agent-specific attributes (gen_ai.conversation.id, gen_ai.tool.*, multi-agent handoffs), see the AI Agent Observability guide.

GenAI Metrics​

Custom metrics for dashboards and alerting:

MetricTypePurpose
gen_ai.client.token.usageHistogramToken consumption by model/agent
gen_ai.client.operation.durationHistogramLLM call latency
gen_ai.client.costCounterCost in USD by model/agent
gen_ai.evaluation.scoreHistogramOutput quality scores
gen_ai.client.error.countCounterErrors by provider/type

Next Steps​

  1. Follow the LLM Observability guide for a complete Python setup walkthrough
  2. Follow the Rust LLM Observability guide for Rust AI applications with manual GenAI instrumentation
  3. Follow the Spring AI LLM Observability guide for Java Spring AI applications with three-layer instrumentation
  4. Set up auto-instrumentation for your web framework if you haven't already
  5. Configure the OpenTelemetry Collector to export telemetry to base14 Scout
Was this page helpful?