Practice Exams:

Building Tool-Using Agents in Microsoft Foundry

 

An agent becomes operationally useful when it can do more than generate language. It may need to search enterprise knowledge, call an API, inspect a database, execute a function, create a ticket, or hand work to another agent. Tools make those capabilities possible, but they also introduce contracts, permissions, error states, latency, and side effects that ordinary chat applications can avoid.

The current AI-103 role explicitly covers agents that integrate retrieval, function calling, memory, APIs, knowledge stores, search, Content Understanding, and custom functions. An Azure AI Apps and Agents Developer therefore needs to think of tool use as software integration mediated by a probabilistic planner, not as a magical extension of the prompt.

Microsoft Foundry provides managed agent capabilities, tools, toolboxes, observability, identity, and deployment options, but good outcomes still depend on the developer making the tool surface understandable, narrow, testable, and safe.

Give each tool one clear reason to exist

A model chooses tools partly from their names, descriptions, schemas, and instructions. If several tools overlap or expose vague capabilities, selection becomes harder. A useful tool should represent a distinct action or source of knowledge, such as “retrieve order status,” “create support note,” or “search policy documents.”

This reflects the broader logic of rational agents: behavior improves when available actions and observations are well defined. A tool catalog full of broad, ambiguous operations gives the agent a noisy action space and makes failures harder to explain.

Design schemas for the model and the service

Tool parameters have to be precise enough for the downstream system and understandable enough for the model to populate correctly. Prefer descriptive field names, bounded enumerations, required-versus-optional clarity, and validation rules that catch malformed input before an external call is made.

Do not expose internal database or API complexity directly unless the agent genuinely needs it. A business-facing function can translate a simple, validated request into several backend calls. That abstraction makes tool selection easier and prevents the model from constructing implementation details that should remain deterministic.

Keep authorization outside the model’s discretion

An agent can decide that a tool is relevant, but it should not decide whether it is authorized. The runtime or downstream service must enforce identity, role, resource scope, and policy. A prompt instruction that says “only managers may approve refunds” is not an access-control mechanism.

This separation becomes more important as tools gain side effects. The architecture described by AB-100 defines which business actions belong in the agent workflow; the AI-103 implementation should enforce those boundaries through identities, permissions, approval steps, and constrained operations.

Use tool-choice controls deliberately

Some requests should always use a tool, some should never use one, and others can leave the decision to the model. Treating every invocation as fully automatic can produce unnecessary calls or allow the agent to answer from memory when a real-time lookup was required.

Foundry agent patterns support explicit control over tool use. The important engineering idea is to match that control to the task. A “check current account balance” intent should require authoritative retrieval, while a conceptual explanation may need no external action. Deterministic routing can coexist with agent reasoning.

Make errors part of the contract

Tools fail because networks time out, services return validation errors, records are missing, users lack permission, or dependencies are unavailable. A good tool interface returns structured error information the agent can interpret without exposing sensitive implementation details.

The agent also needs instructions for what to do next. Some failures justify one retry; others require corrected input, a user clarification, or escalation. Unlimited retry loops can amplify cost and duplicate side effects. Error handling should be designed before production rather than discovered from incident logs.

Protect idempotency when actions can repeat

Model-driven workflows can call the same tool twice because of retries, conversation ambiguity, or orchestration bugs. If the tool creates records, sends messages, charges money, or changes access, a duplicate call may be much worse than a failed call. The service should support idempotency keys, transaction checks, or other safeguards where repeated execution matters.

The agent should also receive enough response detail to recognize that an action already succeeded. Returning only “OK” makes it difficult to distinguish a completed transaction from a generic success message. Stable identifiers and explicit status improve both recovery and auditability.

Observe tool calls as first-class production events

Chat transcripts are not sufficient for debugging tool-using agents. Operators need to see which tool was selected, the sanitized parameters, execution duration, response status, downstream identifiers, and how the result affected the next agent step. Distributed traces are especially valuable when one request crosses retrieval, model, tool, and application boundaries.

Tool telemetry should support quality evaluation as well as incident response. Repeated selection of the wrong tool, excessive retries, or unusual parameter patterns can reveal instruction problems before users report failures. Observability closes the loop between agent behavior and software operations.

Use approvals for consequential actions

Human approval is useful when the agent can prepare an action safely but should not execute it without confirmation. The approval request needs to show the proposed operation, important parameters, and the evidence behind the decision. A vague “approve?” button simply moves uncertainty to the reviewer.

This pattern is particularly valuable during gradual rollout. Teams can begin with all writes requiring approval, collect evidence about accuracy, and later automate narrowly defined low-risk actions. Autonomy can expand as confidence grows instead of being granted on day one.

Tool descriptions should be written for decision quality rather than marketing. Explain the conditions under which a tool is appropriate, what it returns, and any important limitations. If two tools can both retrieve customer data, for example, the descriptions should distinguish whether one is authoritative for billing while the other is optimized for support context. Clear differences reduce arbitrary selection.

Data returned by a tool should be treated as untrusted input even when the service itself is trusted. External systems can contain malformed text, user-generated content, or instructions that were never intended for an agent. The application should preserve the distinction between tool data and system instructions so retrieved content cannot redefine the agent’s authority.

Latency budgeting becomes important as tool chains grow. A user request that triggers a search, two APIs, an approval check, and another model call can become slow even when each dependency performs reasonably. Developers should measure the critical path and consider parallel calls where dependencies allow it, while avoiding concurrency that creates duplicate or conflicting actions.

State management deserves the same care. A conversation may need to remember selected customer, pending approval, completed tool calls, and unresolved errors. That state should be represented explicitly rather than inferred repeatedly from a long transcript. Explicit state makes retries and multi-turn recovery more reliable and reduces the tokens needed to reconstruct what already happened.

Testing should include intentional tool misuse. Ask the agent for an action it lacks permission to perform, provide invalid identifiers, return contradictory tool results, and simulate a stale schema. The expected behavior should be known before launch. A robust agent does not merely succeed when tools cooperate; it fails in predictable, bounded ways when they do not.

Finally, tools should expose the minimum data required for the next reasoning step. Returning an entire customer record when the agent needs only status and renewal date increases privacy exposure and prompt size. Narrow outputs improve both security and model focus, which is another reason to design purpose-built tool contracts rather than giving agents raw access to general APIs.

Contract testing is especially useful when tools are owned by other teams. The agent team should be able to verify that required fields, response shapes, error codes, and authorization behavior remain compatible after a service release. That protects the agent from downstream changes that would otherwise appear as mysterious reasoning failures.

For multi-agent systems, a remote agent should be treated like another tool unless there is a reason to grant broader trust. Define what tasks may be delegated, what context may be forwarded, what identity is used, and how the calling agent verifies the result. Agent-to-agent communication does not remove the need for API-style boundaries; it makes them more important.

Sandbox environments are useful for new tools because they let developers observe selection behavior without exposing production data or side effects. A tool can graduate only after its schema, permission scope, failure behavior, and trace evidence are understood.

Tool ownership should also be visible. Someone needs to approve schema changes, permission changes, and retirement plans, because an agent can depend on a tool long after the original developer has moved on. Clear ownership prevents silent breakage in shared agent platforms.

Version tools and agent instructions together

Tool behavior changes over time. Parameters are added, APIs are replaced, permissions change, and business rules evolve. If the agent instructions describe an old version of a tool, selection and parameter quality can deteriorate even though each component works independently.

Foundry toolboxes and managed agent definitions support a lifecycle mindset: version, test, promote, and observe. The practical objective is to make an agent release reproducible. A known combination of model, instructions, tool schemas, connections, and evaluation tests should be able to move through environments as a controlled unit.

Tool-using agents are strongest when the tool layer behaves like disciplined application architecture. Clear responsibilities, narrow schemas, external authorization, intentional routing, structured errors, idempotency, tracing, approvals, and versioning make agent behavior more reliable. The model remains probabilistic, but the surrounding system can still impose deterministic boundaries on what actions are possible and how they are verified.

Related Posts

• The First 15 Minutes of Incident Triage

• Backups, Recovery, and Continuity Are Different Problems

• Reading an Azure Cost Spike Like an Administrator

• How Azure Subscriptions, Policy, and Locks Work Together

• IPv6 Without the Fear: What Changes and What Stays Familiar

• Identity Is the New Security Perimeter

• Guardrails, Moderation, and the Limits of Model Safety Controls

• Fine-Tuning or Better Retrieval?

• Wireless Design Starts With RF

• Infrastructure as Code for CLI-First Network Teams