Practice Exams:

Microsoft AI-103: Tracing AI Agents in Azure

Agent tracing answers the production question that eventually matters most: why did the agent do that? A final response is not enough when the system retrieves documents, calls models, uses tools, writes memory, or coordinates other agents. Operators need to see the sequence of spans that produced the outcome, how long each step took, which tool arguments were sent, and where an error or latency spike entered the flow.

Microsoft Foundry tracing uses OpenTelemetry conventions and stores trace data in connected Azure Monitor Application Insights. For agents hosted in Foundry, Microsoft recommends starting with server-side tracing after connecting Application Insights. Client-side instrumentation can then extend visibility into custom application code and external frameworks.

Tracing is therefore a production requirement inside Azure AI engineering, not merely a debugging feature.

Start with server-side tracing

Foundry can automatically capture traces for hosted agents once Application Insights is connected. This gives a useful baseline without requiring instrumentation code in the application.

Traces can show conversations, responses, ordered actions, model calls, tool calls, duration, status, and token information depending on the path.

This is often enough to identify whether a problem belongs to the model, a tool, or the managed agent flow before custom telemetry is added.

Add client spans around custom logic

Applications often do work before and after the agent call: authorization, retrieval, validation, queues, or custom tool execution.

Add OpenTelemetry spans around those boundaries so the trace reflects the complete request instead of stopping at the managed agent endpoint.

AI observability is strongest when prompt, retrieval, model, tool, and application spans can be followed in one correlated trace.

Use consistent semantic conventions

OpenTelemetry GenAI conventions give teams a shared vocabulary for model calls, agent invocations, workflows, tools, and memory operations. Foundry also supports multi-agent observability conventions that connect orchestration and child-agent spans.

Consistency makes traces queryable across frameworks. A tool call should be recognizable whether the agent was built with Agent Framework, Semantic Kernel, LangGraph, or another supported runtime.

Use stable attributes for agent version, model deployment, tool name, workflow, and environment.

Correlate tool calls with outcomes

A trace that stops at “tool returned 200” is incomplete if the tool updated the wrong business object. Include safe identifiers and operation status that connect the technical call to the business result.

Tool calling is easier to debug when the trace shows the requested tool, validated arguments, downstream response, and agent continuation.

Avoid putting secrets or unnecessary personal data into span attributes merely to make debugging easier.

Trace memory and retrieval separately

Memory and RAG can both supply context, but they have different trust, freshness, and lifecycle rules. Tracing should reveal whether a response was influenced by a memory read, a retrieval query, or both.

Agent memory operations should appear as distinct steps from search or knowledge retrieval.

This matters when an incorrect answer came from stale memory rather than a bad retrieval result.

Protect trace data like production data

Microsoft notes that Foundry tracing can capture prompts, model inputs and outputs, tool arguments and results, token usage, errors, and other operational telemetry.

That data can be sensitive. Apply access control, retention, redaction, and minimization. Do not put credentials or tokens into prompts, tool arguments, or custom span attributes.

GenAI observability should improve visibility without creating an ungoverned copy of sensitive application data.

Use traces to investigate latency

An agent can be slow because of generation, retrieval, tool calls, queueing, network latency, or repeated planning steps. End-to-end duration alone does not identify the bottleneck.

Span durations show where time was spent. Latency tuning should use that evidence before changing model size, deployment type, or concurrency.

Look at percentiles and slow traces, not only averages, so intermittent tool and network problems remain visible.

Connect traces to evaluation

Quality and observability are more powerful together. Foundry can evaluate production traces and convert selected trace examples into reusable datasets.

Online evaluation should carry trace or evaluation-run identifiers so a low quality score can be opened against the execution path that produced it.

Important incidents can then become regression tests instead of one-off debugging sessions.

Trace only what helps operation

More telemetry is not always better. High-volume prompt bodies, full tool payloads, and long retention can increase cost and privacy risk.

AI metrics should guide what needs full tracing, what can be sampled, and which attributes are necessary for diagnosis.

Review trace volume after new tools or agents launch. A workflow that fans out into many calls can multiply telemetry unexpectedly, so sampling and retention may need to change with architecture.

For current AI-103 work, the durable tracing model is to start with server-side Foundry traces, extend them with OpenTelemetry around custom code, correlate tools and memory, minimize sensitive data, and connect operational evidence to evaluation and release decisions.

Trace design should preserve parent-child relationships across asynchronous work. A user request may enqueue a job, trigger a function, invoke an agent later, and then call several tools. Pass trace context or a stable correlation identifier across those boundaries so the final execution can still be connected to the initiating request. Without that propagation, long-running agent workflows appear as unrelated telemetry fragments.

Sampling strategy matters at scale. Full tracing is valuable during development and incidents, but high-volume production workloads may need probabilistic or rule-based sampling. Keep errors, high-latency requests, denied actions, and risky tool paths at a higher sampling rate than routine successful interactions. Sampling should preserve enough normal traffic to provide a baseline rather than storing only failures.

Use traces to understand agent loops. Repeated planning steps, tool retries, or agent-to-agent handoffs can consume tokens without improving the outcome. A trace can reveal that the system is technically successful but operationally inefficient. Add limits on tool attempts, workflow hops, or repeated memory lookups when traces show patterns that are expensive or unstable.

Release comparisons become easier when traces include version metadata. Record the model deployment, prompt or agent version, toolbox version, application build, and retrieval index where possible. During canary rollout, this allows operators to compare candidate and baseline latency, errors, token use, and tool behavior on similar traffic rather than relying on one aggregate dashboard.

Tracing should feed runbooks. Define how support engineers move from a user-reported issue to a trace, which spans identify model, retrieval, and tool work, what data is safe to inspect, and when the case should move to a security or privacy team. Observability creates value when it shortens diagnosis and improves decisions, not simply when the portal contains more telemetry.

Trace queries should be designed before an incident. Save useful Application Insights or Log Analytics queries for high latency, failed tool calls, repeated agent loops, large token consumption, and specific agent versions. Operators work faster when the common diagnostic questions already have tested queries rather than being assembled under pressure.

Retention should match investigation needs and policy. Short retention may reduce cost but make slow-moving quality problems difficult to study; excessive retention can increase privacy and compliance risk. Different telemetry classes may justify different retention periods, especially when prompts or tool payloads contain sensitive content.

Cross-team access should also be deliberate. Developers may need broad trace detail in nonproduction, while production access can be restricted to support or operations roles. Foundry and Application Insights permissions should reflect the sensitivity of the captured data rather than assuming every project contributor needs to read every production interaction.

Distributed traces should also survive queues and background work. Carry W3C trace context or a stable correlation identifier through Service Bus, Functions, Durable Task, and custom workers so an asynchronous tool or evaluation job can still be connected to the original request. This is especially important for long-running agents where the user-facing response may arrive long before every follow-up action completes.

Use trace attributes sparingly for high-cardinality values. User IDs, full URLs, document IDs, and raw prompts can create cost, privacy, and query-performance problems. Prefer stable categorical fields for dashboards and keep detailed identifiers only where they are genuinely needed for investigation.

Tracing should be tested like any other dependency. After changing frameworks, agents, toolboxes, or SDK versions, run a known workflow and confirm that expected spans still appear. Missing telemetry is easiest to fix before the first incident that depends on it.

During multi-agent investigations, preserve the agent name and parent workflow on each span so one child agent’s delay is not mistaken for the orchestrator. This makes handoffs and recursive calls understandable when several agents participate in one user request.

Related Posts

• Mastering the Azure AI Engineer (AI-102) Exam

• How to Ace the Microsoft Certified Azure AI Engineer Exam (AI-102)

• Microsoft AI-300: Monitor Model Drift Without Chasing Noise

• Microsoft AI-300: Machine Learning CI/CD Needs More Than a Build Pipeline

• Microsoft AI-103: Computer Vision and Multimodal AI in One Application

• Microsoft Business AI Systems

• Microsoft AI-103: Azure AI Content Safety in Practice

• Microsoft AI-103: Deploying Fine-Tuned Models on Azure

• Microsoft AI-103: Managing Agent Memory on Azure

• Microsoft AI-103: Reranking for Better Azure RAG