Skip to main content

Microsoft Agent Framework

Microsoft Agent Framework has OpenTelemetry built in. With instrumentation on, which is the default, every agent run emits invoke_agent, chat and execute_tool spans, gen_ai.* metrics, and message events when sensitive data is on. This page covers the Python packages.

TL;DR

Call configure_otel_providers() once at startup, or set your own providers and call enable_instrumentation(). Pass enable_sensitive_data=True to record messages, tool arguments and tool results. The Ollama client leaves server.address as Unknown and sets no conversation ID, so add those in a span processor.

Note: For framework-agnostic agent patterns, see AI Agent Observability. For other Python agent frameworks, see Strands Agents, Google ADK and OpenAI Agents SDK.

Running this in production

Storing and querying these traces at production volume is what base14 Scout does. Check out Scout LLM Observability.

Who This Guide Is For​

  • Python developers running Microsoft Agent Framework who want traces, metrics and message events in an OpenTelemetry backend.
  • Teams running it on a local model through agent-framework-ollama.
  • Teams nesting one agent inside another with as_tool, who want one trace per request across both.

Overview​

  • Turn the built-in instrumentation on and send it to a collector.
  • Read the span tree of a run, including an agent called as a tool.
  • Put a conversation ID and request IDs on spans with a span processor.
  • Turn content capture on with sensitive data.
  • Recognize a budget stop and a failed tool in a trace.
  • Add cost and the real server in a span exporter.

Signals​

SignalWhat Agent Framework emitsWhat the example adds
Tracesinvoke_agent, chat and execute_tool spans.Request IDs, the conversation ID, prompt versions and the server port through a span processor. Cost in a span exporter. FastAPI, httpx and psycopg spans, and one hand-written span.
Metricsgen_ai.client.operation.duration, gen_ai.client.token.usage and agent_framework.function.invocation.duration.Application counters and a duration histogram under base14.filing.*. http.server.* and http.client.duration.
LogsMessage events for each model call, with sensitive data on.The OpenTelemetry LoggingHandler on the root logger, so every line carries the trace and span ID.

Prerequisites​

  • Python 3.10 or later. The example uses 3.14.
  • A model with tool calling. The Quick Start and the example use Ollama with qwen3.5:9B.
  • An OpenTelemetry Collector. See Docker Compose Setup.
  • A base14 Scout account, optional. The example runs without one.

Compatibility Matrix​

ComponentVersion in the example
agent-framework-core1.19.0
agent-framework-ollama1.0.0b260813, which pins ollama 0.5.3
opentelemetry-sdk, opentelemetry-api1.45.0
opentelemetry-exporter-otlp-proto-http1.45.0
opentelemetry-instrumentation-fastapi, -httpx, -psycopg, -logging0.66b0
fastapi0.141.1
Python3.14
Ollama0.34.2, with qwen3.5:9B and gemma4:e2b
OpenTelemetry Collector Contrib0.161.0
Exampleai-filing-analyst with FILING_FRAMEWORK=maf

Last verified 2026-09-29 with Agent Framework 1.19.0. The GenAI conventions are in Development status and agent-framework-ollama is a beta, so attribute names can change between releases. Pin exact versions and re-check the spans after each upgrade.

Installation​

Terminal
uv add agent-framework-core==1.19.0 \
agent-framework-ollama==1.0.0b260813 \
opentelemetry-sdk==1.45.0 \
opentelemetry-exporter-otlp-proto-http==1.45.0

agent-framework-core depends on the OpenTelemetry API only. Install the SDK and an OTLP exporter to send anything. configure_otel_providers() uses gRPC unless OTEL_EXPORTER_OTLP_PROTOCOL says otherwise. For gRPC, install opentelemetry-exporter-otlp-proto-grpc instead.

Quick Start​

This file is a minimal starting point, not part of the example. It sets up tracing, metrics and logs, turns sensitive data on, and runs one agent with one tool on Ollama. Save it as quickstart.py:

quickstart.py
import asyncio

from agent_framework import Agent, tool
from agent_framework.observability import configure_otel_providers
from agent_framework_ollama import OllamaChatClient
from opentelemetry import metrics, trace

configure_otel_providers(enable_sensitive_data=True)


@tool(approval_mode="never_require")
def order_status(order_id: str) -> str:
"""Return the shipping status of an order."""
return "shipped" if order_id == "A-100" else "not found"


agent = Agent(
OllamaChatClient(host="http://localhost:11434", model="qwen3.5:9B"),
"Look up the order with the order_status tool, then answer in one sentence.",
name="order-agent",
tools=[order_status],
)


async def main() -> None:
result = await agent.run("Where is order A-100?")
print(result.text)


asyncio.run(main())
trace.get_tracer_provider().shutdown()
metrics.get_meter_provider().shutdown()

Pull the model, point the file at your collector and run it:

Terminal
ollama pull qwen3.5:9B
export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318
export OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
export OTEL_SERVICE_NAME=order-agent
python quickstart.py

It prints the answer, for example Order A-100 has been shipped. Your backend then shows one trace for the run under the service order-agent:

The Quick Start trace
invoke_agent order-agent
|-- chat qwen3.5:9B asks for the tool call
|-- execute_tool order_status
`-- chat qwen3.5:9B writes the answer

The three metrics and the message events arrive with the same service name. The explicit shutdowns flush the metrics before the process exits. If no trace shows up, check the endpoint and protocol and see Troubleshooting.

configure_otel_providers() sets service.version to the Agent Framework version unless you pass service_version or set OTEL_SERVICE_VERSION.

Configuration​

configure_otel_providers() creates the providers and exporters from the standard OTEL_* variables and turns instrumentation on. Call it once. To keep providers you already set, call enable_instrumentation() instead. The example sets its providers, with its own exporter wrapper and logging handler, in telemetry.py, then:

frameworks/maf.py (condensed)
def from_settings(settings: Settings) -> MafFramework:
capture = os.environ.get("OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT", "true").strip().lower() != "false"
enable_instrumentation(enable_sensitive_data=capture)
return MafFramework(ollama_clients(settings.ollama_base_url), settings.ollama_think)

Environment Variables​

VariableValue in the exampleRead by
OTEL_SERVICE_NAMEai-filing-analystThe SDK resource.
OTEL_EXPORTER_OTLP_ENDPOINThttp://otel-collector:4318The OTLP exporters.
OTEL_EXPORTER_OTLP_PROTOCOLhttp/protobufThe OTLP exporters. configure_otel_providers() defaults to grpc.
OTEL_SEMCONV_STABILITY_OPT_INgen_ai_latest_experimentalAgent Framework. Selects the latest GenAI conventions, which is also its default. A value without that token selects the v1.36.0 conventions, without tool arguments and results on spans.
ENABLE_INSTRUMENTATIONNot set.Agent Framework. Default true.
ENABLE_SENSITIVE_DATANot set. The example passes enable_sensitive_data.Agent Framework. Default false.
ENABLE_MESSAGE_EVENTSNot set.Agent Framework. Default true. Emits message events when sensitive data is on.
OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENTtrueThe example, which maps it to enable_sensitive_data. Agent Framework does not read it.
OTEL_ATTRIBUTE_VALUE_LENGTH_LIMIT4096The SDK. Caps each captured value.

What Agent Framework Emits​

This is the ranking question from the example, one trace per HTTP request. The psycopg SELECT and INSERT spans and the ASGI http receive and http send spans are left out.

One ranking question
POST /questions
|-- invoke_agent analyst
| |-- chat qwen3.5:9B
| |-- execute_tool rank_among_filers
| | `-- invoke_agent ranking
| | |-- chat gemma4:e2b
| | |-- execute_tool frame_values
| | | `-- GET SEC frames API
| | `-- chat gemma4:e2b
| |-- chat qwen3.5:9B
| `-- execute_tool FilingAnswer the typed answer
`-- filing.verify_answer hand-written

Agent Framework emits one invoke_agent <agent> per agent.run, one chat <model> per model call and one execute_tool <tool> per tool call. An agent attached with as_tool runs inside its execute_tool span, so both agents share one trace. When the analyst ends a run in text without calling FilingAnswer, the example runs it once more in the same session with a reminder, which adds a second invoke_agent analyst.

The attributes worth knowing:

  • gen_ai.agent.name, gen_ai.agent.id and gen_ai.request.model on invoke_agent, with the run's total gen_ai.usage.input_tokens and gen_ai.usage.output_tokens.
  • gen_ai.request.model, gen_ai.response.model, gen_ai.response.finish_reasons and the call's token counts on chat.
  • gen_ai.tool.name, gen_ai.tool.call.id, gen_ai.tool.type and gen_ai.tool.description on execute_tool.
  • With sensitive data on, gen_ai.system_instructions, gen_ai.input.messages and gen_ai.output.messages on invoke_agent and chat, gen_ai.tool.definitions on invoke_agent, and gen_ai.tool.call.arguments and gen_ai.tool.call.result on execute_tool.

gen_ai.provider.name is ollama on chat and microsoft.agent_framework on invoke_agent. The Ollama client sets server.address to Unknown and no server.port.

Agent Framework Metrics​

MetricWhat it measuresAttributes
gen_ai.client.operation.durationModel call duration.gen_ai.operation.name, gen_ai.provider.name, gen_ai.request.model, gen_ai.response.model, server.address, and error.type on a failed call.
gen_ai.client.token.usageTokens per model call.The same, without error.type, plus gen_ai.token.type.
agent_framework.function.invocation.durationTool call duration.The execute_tool span's attributes, agent_framework.function.name, and error.type on a failed call.

There is no agent-level duration metric. invoke_agent spans carry the run duration.

agent_framework.function.invocation.duration takes the tool span's attributes, which include gen_ai.tool.call.id on every call and gen_ai.tool.call.arguments with sensitive data on. Each tool call then makes a new time series. Drop those attributes with a metric view or in the collector before they reach a metrics backend:

Keep tool metrics to the tool name
from opentelemetry.sdk.metrics.view import View

tool_duration = View(
instrument_name="agent_framework.function.invocation.duration",
attribute_keys={"gen_ai.operation.name", "gen_ai.tool.name", "gen_ai.tool.type", "error.type"},
)

Pass the view in MeterProvider(views=[...]), or in configure_otel_providers(views=[...]).

Message Events​

With sensitive data on, Agent Framework also emits the v1.36.0 GenAI message events as log records under the scope agent_framework: gen_ai.system.message, gen_ai.user.message, gen_ai.assistant.message, gen_ai.tool.message and gen_ai.choice. Each carries the trace and span ID of its chat span and gen_ai.system=ollama. Set ENABLE_MESSAGE_EVENTS=false, or pass enable_message_events=False, to keep content on the spans only. They need a global logger provider, which configure_otel_providers() sets.

Request Attributes​

Agent Framework sets gen_ai.conversation.id only from a session ID that the model service manages. A local Ollama session has none. The example sets the question's attributes in a context variable around the run, and a span processor copies them onto each GenAI span as it starts:

telemetry.py (condensed)
class AgentRunAttributesProcessor(SpanProcessor):
def on_start(self, span: Span, parent_context: Context | None = None) -> None:
run = _agent_run.get()
if run is None or not span.name.startswith(GEN_AI_SPAN_PREFIXES):
return
attributes = span.attributes or {}
added = {**run.question, **_agent_or_model(run, span.name, attributes)}
span.set_attributes({key: value for key, value in added.items() if key not in attributes})

run.question holds the request ID, gen_ai.conversation.id and the ticker. The processor also adds server.port from the Ollama URL, and the prompt version and model digest from the span's agent name or model. It keeps any attribute the framework already set, so server.address stays Unknown. To replace it, overwrite it in the processor instead.

Agents as Tools​

agent.as_tool(name=..., description=...) wraps an agent as a tool of another agent:

frameworks/maf.py (condensed)
ranking = Agent(
clients(config.ranking_model),
config.ranking_prompt.system,
name="ranking",
tools=maf_tools(tools.ranking),
default_options=options,
middleware=[middleware.chat(analyst=False), middleware.function()],
)
analyst = Agent(
clients(config.analyst_model),
analyst_instructions(config, FINISH_RULE),
name="analyst",
tools=[
*maf_tools(tools.analyst),
*maf_tools([answer_tool(sink)]),
ranking.as_tool(name="rank_among_filers", description=RANKING_TOOL_DESCRIPTION),
],
default_options=options,
middleware=[middleware.chat(analyst=True), middleware.function()],
)

A tool that fails inside the inner agent shows on its execute_tool span with error status. The outer execute_tool rank_among_filers span ends without error. The example's function middleware rewrites the ranking tool's result: it appends the frame's facts to the inner agent's reply, or replaces the reply with a fixed line when the frames fetch failed.

Structured Output on Ollama​

The response_format option asks the model for JSON in a schema. The Ollama client sends it as format on every call of the run, so the model can no longer call a tool. The example gives the analyst a FilingAnswer function tool instead, whose parameters are the answer's fields. Function middleware ends the run once it is called:

frameworks/maf.py (condensed)
if context.function.name == "FilingAnswer":
await call_next()
raise MiddlewareTermination

The typed answer shows as execute_tool FilingAnswer.

Budgets​

Chat middleware runs before each model call and function middleware before each tool call. Registered on both agents, they see every call in a run, including those of an agent called as a tool. The example counts both with a plain CallBudget counter, which raises past the budget:

frameworks/maf.py (condensed)
@chat_middleware
async def on_chat(context: ChatContext, call_next: Next) -> None:
budget.count_model() # raises past the budget
await call_next()


@function_middleware
async def on_function(context: FunctionInvocationContext, call_next: Next) -> None:
exceeded = budget.count_tool()
if exceeded is not None:
raise exceeded
await call_next()

The exception ends invoke_agent with error status and error.type set to the exception's class name, BudgetExceeded. Agent Framework records the exception on the span, so the type needs no exporter.

Agent Framework has no wall-clock limit. The example wraps the run in asyncio.wait_for, which cancels it at the deadline. A cancelled run does not mark invoke_agent as an error. The example's server span ends with error status, a 504 and base14.filing.outcome=timeout.

Logs and Trace Correlation​

Apart from the message events, Agent Framework writes to Python logging only. The example adds the OpenTelemetry LoggingHandler to the root logger, as in Strands Agents, so every application log record, including Agent Framework's own, goes to the collector with the trace ID and span ID of the span it was written under.

Content Capture​

Content capture follows sensitive data, off by default. With enable_sensitive_data=True or ENABLE_SENSITIVE_DATA=true, Agent Framework records messages, system instructions and tool definitions on the spans, tool arguments and results on execute_tool, and the message events. Agent Framework does not read OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT. The example maps it to enable_sensitive_data so all its frameworks follow one setting.

Adding Cost and Error Type with a Span Exporter​

The example wraps the OTLP span exporter to add, on the way out:

  • base14.gen_ai.cost on each chat span, from the token counts and a price table. A model with no price, such as a local Ollama model, gets 0 and base14.gen_ai.cost.simulated=true.
  • error.type on a failed span that has none, from the recorded exception or the HTTP status. Agent Framework's own spans already have it.

gen_ai.provider.name on chat is already ollama. See CostAndErrorAttributingSpanExporter in the example's telemetry.py.

Known Gaps​

As of Agent Framework 1.19.0 and agent-framework-ollama 1.0.0b260813, verified 2026-09-29:

  • server.address is Unknown on Ollama chat spans and metrics, and there is no server.port. Overwrite them in a span processor.
  • No conversation ID with a local session. Set gen_ai.conversation.id in a span processor.
  • The tool duration metric carries gen_ai.tool.call.id, and the tool arguments with sensitive data on. Drop them with a metric view.
  • invoke_agent reports the provider as microsoft.agent_framework. The model's provider is on chat.
  • No agent-level duration metric.
  • service.version defaults to the Agent Framework version with configure_otel_providers(), even when OTEL_RESOURCE_ATTRIBUTES sets one. Pass service_version or set OTEL_SERVICE_VERSION.
  • response_format cannot be used with tools on Ollama. Ollama applies the schema to every call. Use a function tool and end the run in middleware.
  • agent-framework-ollama is a beta. It pins ollama below 0.5.4.
  • Message events use the v1.36.0 names and gen_ai.system, while the spans use the latest conventions and gen_ai.provider.name.

What to Look For in Scout​

Follow one request across both agents​

Search spans by gen_ai.conversation.id or your own request ID attribute. The trace shows invoke_agent ranking under execute_tool rank_among_filers, with its own chat spans.

Find failed runs and why​

Filter spans on status = Error and group by error.type. BudgetExceeded on invoke_agent marks a budget stop. A model server that cannot be reached shows as ChatClientException on chat and invoke_agent. A timeout leaves invoke_agent without error status. The example's server span ends with error status, a 504 and base14.filing.outcome=timeout.

Find a tool that failed inside a run that completed​

Filter execute_tool spans on error status. A tool that raises gets error.type set to the exception's class name, and the error goes back to the model, so the run can still complete. A failed frames fetch in the ranking agent shows here as SecUnavailable.

Read tokens by model​

Sum gen_ai.usage.input_tokens and gen_ai.usage.output_tokens on chat spans, grouped by gen_ai.request.model, or read the gen_ai.client.token.usage metric by gen_ai.request.model and gen_ai.token.type. The totals on invoke_agent repeat the chat counts, so do not add both.

Production Patterns​

  • Leave sensitive data off for real data. It is off by default.
  • Cap attribute length. Tool results can be long. The example sets OTEL_ATTRIBUTE_VALUE_LENGTH_LIMIT=4096.
  • Keep the tool metric's attributes small with a view, as in Agent Framework Metrics.
  • Bound every run twice. A call budget in middleware stops a loop, and asyncio.wait_for stops a slow or stalled model or tool.
  • Put the prompt version and model digest on spans. An Ollama tag can point at new weights. The example reads each model's digest from Ollama at startup.
  • Send through a collector. The example exports OTLP HTTP to a collector, which authenticates to Scout and keeps a debug exporter for local checks.
  • Keep fault injection off. The example's fault fields are refused unless FILING_FAULTS_ENABLED=true. Set it only for the scenario harness.

Running Your Application​

Run quickstart.py as in Quick Start. For the full example on Agent Framework:

Terminal
git clone https://github.com/base-14/examples.git
cd examples/python/ai-filing-analyst
cp .env.example .env
ollama pull qwen3.5:9B
ollama pull gemma4:e2b
make docker-up FRAMEWORK=maf

Set SEC_USER_AGENT in .env to your organisation's name and a contact email first. Then ask a question:

Terminal
curl -s -X POST http://localhost:8000/questions \
-H 'Content-Type: application/json' \
-d '{"ticker": "KVYO", "question": "What was Klaviyo'"'"'s revenue for its latest fiscal year?"}' | jq

scripts/test-api.sh runs seventeen scenarios, eight with injected faults, and scripts/verify-scout.sh checks the telemetry each one produced. The fault scenarios need the stack started with FILING_FAULTS_ENABLED=true make docker-up FRAMEWORK=maf.

Troubleshooting​

No agent spans​

Nothing set a tracer provider. Call configure_otel_providers(), or set your own and call enable_instrumentation(). If ENABLE_INSTRUMENTATION=false or disable_instrumentation() was called, pass force=True.

Spans but no metrics​

The process exited before the periodic metric export. Shut the meter provider down before exit, as the Quick Start does.

Nothing arrives at the collector​

configure_otel_providers() defaults to gRPC. For port 4318, set OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf and install the HTTP exporter.

No messages on spans​

Sensitive data is off. Pass enable_sensitive_data=True or set ENABLE_SENSITIVE_DATA=true.

No tool arguments or results on execute_tool​

OTEL_SEMCONV_STABILITY_OPT_IN is set without gen_ai_latest_experimental, which selects the v1.36.0 conventions. Add the token.

FAQ​

Does Microsoft Agent Framework support OpenTelemetry?​

Yes, the Agent Framework Python packages have OpenTelemetry built in and on by default. configure_otel_providers() sets up the exporters, or you can bring your own providers.

Which spans does an Agent Framework run produce?​

One invoke_agent <agent> per run, one chat <model> per model call and one execute_tool <tool> per tool call.

How do I turn on prompt capture in Agent Framework?​

Pass enable_sensitive_data=True to configure_otel_providers() or enable_instrumentation(), or set ENABLE_SENSITIVE_DATA=true.

How do I set the conversation ID in Agent Framework?​

Agent Framework sets gen_ai.conversation.id only from a service-managed session ID. With a local model, add it in a span processor, as in Request Attributes.

Does Agent Framework record cost?​

No, Agent Framework records token counts but not cost. Add a cost attribute in a span exporter, as in Adding Cost and Error Type.

How do I trace one Agent Framework agent calling another?​

Attach the inner agent with as_tool. It runs inside the outer agent's execute_tool span, so both share one trace.

What's Next?​

Scout Platform Features​

Complete Example​

ai-filing-analyst answers questions about a US-listed company's reported financials from SEC XBRL data. An analyst agent on local Ollama calls tools and, for a ranking question, a second agent attached as a tool. A verifier checks every figure against the tool results before the answer is served. FILING_FRAMEWORK=maf runs it on Agent Framework.

ai-filing-analyst/
|-- prompts/ analyst and ranking prompts, named by UTC timestamp
|-- scripts/
| |-- test-api.sh seventeen scenarios
| `-- verify-scout.sh checks the run's telemetry in the collector output
`-- src/filing_analyst/
|-- telemetry.py providers, logging, run attributes, cost and error attributes
|-- frameworks/maf.py the two agents, middleware, answer tool
|-- agents.py what every framework adapter shares
|-- budget.py the call budget counter
|-- tools.py query_facts, compute_ratio, frame_values
`-- verifier.py the grounding check

Source: python/ai-filing-analyst.

References​

Was this page helpful?