Microsoft AI-103: Tool Calling in Azure AI Agents
Tool calling turns an Azure AI agent from a conversational model into a system that can retrieve live information, call APIs, run business operations, or interact with external services. Microsoft Foundry supports custom function tools, OpenAPI tools, MCP servers, built-in tools, and reusable toolboxes. The architecture question is not how many tools an agent can reach; it is how to make each tool understandable, permissioned, testable, and safe to invoke.
Current Foundry function calling follows a clear loop: define a function schema, let the model request a call, execute the function in trusted application code, return the output, and let the agent continue. Toolboxes add centralized authentication, governance, versioning, and MCP-compatible access across agents. Approval controls can require runtime confirmation for selected calls.
This makes tool design a central part of Azure AI engineering.
Give each tool one responsibility
A tool should describe one meaningful capability with a precise name, purpose, and parameter schema. Generic “execute anything” or “call any URL” tools create an enormous attack and reliability surface.
Agent scope should be defined before tool access. If the agent’s job is narrow, the tool set can be narrow too.
Separate read and write operations when their risk differs. A lookup tool and a record-update tool should not be hidden behind one broad function merely because they use the same backend.
Design schemas for models and validators
Tool descriptions help the model decide when and how to call a function. Typed parameters also give the application a contract it can validate before execution.
Do not trust model-generated arguments simply because they match JSON syntax. Check identifiers, ranges, destinations, enum values, and business rules server-side.
Tool control is strongest when the model proposes an action and trusted code decides whether the proposal is valid.
Use tool choice deliberately
Foundry supports control over whether the model may choose tools automatically, must use a tool, or should avoid tool calls for a turn. The exact control surface depends on the runtime, but the design principle is stable.
Use automatic selection when the model genuinely needs judgment. Require a tool when the product contract demands a lookup or validation step. Disable tools when the turn should remain conversational.
The goal is to reduce ambiguous freedom, not to force every turn through a tool.
Authentication should match the actor
MCP and other tool connections can use key-based access, managed identity, project identity, or user-delegated OAuth depending on the scenario.
Agent identity fits when the tool should act as the agent. User delegation is stronger when the tool must respect the signed-in user’s own permissions.
Do not use one broad project credential merely because it simplifies setup. Authorization should reflect who or what is performing the action.
Approvals must be enforced by runtime
Foundry toolboxes and MCP integrations can carry approval requirements. For high-impact actions, the runtime needs to pause the exact pending call, show the proposed tool and arguments, collect approval, and then continue or reject the call.
A prompt sentence saying “ask before acting” is not the same control. Microsoft documentation explicitly notes that approval has to be enforced by the runtime.
Human oversight is effective when workflow state blocks execution until approval exists.
Choose the tool integration by workload
Function calling is useful for application-owned code. OpenAPI tools expose existing HTTP APIs through a formal specification. MCP provides a reusable standard protocol. Azure Functions can host MCP servers or asynchronous queue-based tools.
Toolboxes are valuable when several agents share governed capabilities and a central owner needs authentication policies and versioning.
Foundry tools should therefore be selected by ownership, latency, reuse, and governance rather than novelty.
Handle tool failure as a normal path
Tools time out, return no data, reject authorization, or produce invalid responses. Agents need explicit behavior for those cases.
Bound retries, preserve error categories, and avoid inventing results when a tool fails. A safe response can explain that the data is unavailable or ask for clarification.
Agent tracing should make failed tool calls visible so operators can separate model mistakes from dependency failures.
Treat tool output as untrusted data
Tool results can contain stale data, malformed values, sensitive information, or malicious text. Validate structured outputs and treat free-form text as data rather than instructions.
This matters in multi-agent and RAG systems where one service’s output becomes another model’s context.
Prompt injection defenses should include tool output because a compromised external system can become an indirect instruction channel.
Test tool use as behavior
A good agent test checks whether the right tool was called, with valid arguments, at the right time, and whether the final task succeeded. Natural-language quality alone is insufficient.
Online evaluation and offline agent tests can compare tool choice, task completion, unnecessary calls, latency, and cost across versions.
Keep tool definitions versioned. A changed description or schema can alter model behavior even when the backend implementation is unchanged. Test the new tool version with representative prompts before promoting it to shared use.
For current Azure AI certification work, the durable pattern is clear: narrow capabilities, typed schemas, least-privilege identity, runtime-enforced approval, validated output, explicit failure behavior, and evaluation that measures the action as well as the conversation.
Latency should influence tool design. A slow synchronous MCP or function call becomes conversation latency because the agent often waits for the result before continuing. Long-running work may be better exposed through an asynchronous queue pattern or a tool that returns a job identifier. The tool contract should match the timing characteristics of the underlying operation instead of forcing every integration into one synchronous shape.
Toolboxes can improve governance when many agents need the same capability. Centralized authentication and versioning reduce duplicated configuration, but shared tools also create shared risk. Changes to a default toolbox version can affect several agents at once, so teams should test a new version against representative agent workloads before promoting it. Consumers should be able to identify which toolbox version was active in a trace.
Tool descriptions deserve version control because small wording changes can alter model selection behavior. A description that is too vague can trigger unnecessary calls; a description that is too broad can encourage the model to use a tool outside its intended scope. Keep names and descriptions concise, describe when the tool should and should not be used, and test ambiguous prompts where several tools could appear plausible.
Return values should be designed for models and applications, not only for human readers. Structured results with explicit status, identifiers, units, and error fields are easier to validate than prose. If the backend returns a large payload, transform it into the smallest trustworthy result the agent needs. This reduces token cost and decreases the chance that irrelevant or malicious text influences the next reasoning step.
Operational ownership should be clear for every tool. Someone must own availability, schema changes, access policy, incident response, and deprecation. An agent team should not discover during an outage that the tool is maintained by another group with different SLAs and no shared escalation path. Treat tools as production dependencies with contracts, not as invisible extensions of the model.
Approval design should include what the reviewer sees. A useful approval screen names the tool, target resource, important arguments, expected side effect, and the user or agent requesting it. Asking a person to approve an opaque JSON blob creates the appearance of oversight without giving them enough context to make a meaningful decision.
Tool deprecation needs a migration plan. Removing or renaming a tool can break prompts, agent versions, and workflows that still reference it. Keep old versions available through a defined sunset window, update consumers deliberately, and use traces to confirm that no active agent is still calling the retired capability.
Finally, tools should not become a hidden way around normal architecture standards. The backend still needs authentication, input validation, logging, rate limits, and service ownership. An MCP or function schema makes the capability discoverable to the model; it does not replace the security and reliability practices that would apply to the same API outside AI.
Observe call frequency by tool as well. A sudden rise can indicate a prompt regression, routing bug, attack, or changed user behavior. Tool analytics should show success rate, latency, denial rate, approval rate, and cost so the team can distinguish a useful capability from one that is merely being invoked often.
Keep a small fallback path for critical tools so temporary discovery or toolbox failures do not force the agent to invent an answer.