Practice Exams:

Anthropic CCDV-F: Claude Tool-Calling Patterns

Tool calling is easiest to reason about when it is treated as a protocol rather than as magic. Claude receives a set of tool definitions, decides that one or more operations are useful, and emits structured calls. Your application executes client-side tools and returns structured results; server-side tools can be executed by Anthropic’s infrastructure. The agent loop continues until the model has enough information to answer or stops for another reason.

Within Claude Development, the best pattern depends on side effects, latency, independence between calls, and the amount of control the application needs. A read-only lookup, a multi-system workflow, and a production deployment should not use the same execution policy merely because all three are “tools.”

Use a single read tool for bounded lookup

The simplest pattern is one narrow read tool: look up an order, fetch a record, search documentation, or calculate a deterministic value. The schema should describe exactly what can be requested, and the result should be compact enough for the model to reason over without receiving an entire database row or raw API response.

This pattern is ideal when the model needs authoritative data that is not in its prompt. The application remains in control of access and can enforce row-level security or tenant boundaries before returning anything. If the lookup fails, return a clear error result rather than fabricating an empty success.

Good read tools also make caching possible. When the same normalized query appears repeatedly and the source allows caching, the executor can reuse a result without changing the model-facing contract. That can reduce latency while preserving deterministic application behavior.

Separate planning from high-impact writes

For actions that create, modify, or delete important data, use a two-stage pattern: let Claude prepare the proposed action, then require policy or human approval before execution. The planning step can show the target, parameters, expected effect, and rollback information. The write tool runs only after the approval condition is satisfied.

This avoids confirmation fatigue because approval is reserved for consequential boundaries. Reading a service status may not need a prompt; deleting an environment should. The article on tool safety explains why the enforcement point belongs in trusted application code rather than in a model instruction that says “ask first.”

A write tool should also be designed for retries. If the operation cannot be safely repeated, include an idempotency key, expected version, or explicit transaction identifier. Returning the created resource ID helps the model and the application reason about what actually happened.

Run independent calls in parallel

If Claude requests independent information from several sources, parallel execution can reduce latency. Examples include fetching account status and documentation, checking two regions, or reading several unrelated configuration records. The executor should preserve each tool-use identifier so results can be matched correctly when they are returned.

Parallelism is inappropriate when one call depends on the output of another or when concurrent writes could conflict. The model may be able to request multiple tools in one turn, but your dispatcher still decides how safely to execute them. Dependency-aware scheduling is an application concern.

Set per-tool concurrency limits. A model that can request ten searches at once should not accidentally overwhelm a legacy backend that supports only two safe concurrent queries. The tool layer is where AI flexibility meets ordinary capacity engineering.

Use a manual loop when you need control

A manual tool loop reads `tool_use` blocks, dispatches them, formats `tool_result` blocks, and continues while the model requests more tools. It is more code than an SDK runner, but it creates explicit checkpoints for authorization, custom tracing, approvals, rate limiting, and application-specific branching.

The design in SDK patterns works well here: centralize the loop in one orchestration layer instead of rebuilding it in every feature. Tool implementations remain ordinary services that can be unit tested without a model, while the orchestrator tests protocol behavior with synthetic calls and results.

Always return a result for every client tool request the model made. Missing or mismatched result identifiers break the conversation protocol. Keep tool results before unrelated user text when continuing a tool turn, and preserve the model’s previous response as required by the API semantics.

Use runners when the standard loop is enough

Anthropic SDK tool runners can automate dispatch, conversation continuation, validation, and error wrapping for supported languages. They are a strong fit when you have a standard agentic loop and would otherwise write plumbing that adds no product-specific value.

Choose the manual loop when you need to inspect tool results before they reach Claude, stop on a specific failure, collect custom audit events, or route approvals through another system. The runner is not “more agentic”; it is simply a higher-level implementation of the same protocol.

Avoid mixing abstractions unpredictably. If one part of the application uses a runner and another manually appends tool messages to the same conversation object, debugging state becomes difficult. Pick one owner for the loop and make extension points explicit.

Return errors that help the next decision

A tool error should say what failed and whether an alternative is possible. “Rate limit exceeded; retry after 30 seconds” gives the model a useful option. “Customer ID not found” may prompt clarification. “Permission denied” should not expose secret policy internals, but it should distinguish lack of access from a nonexistent record.

Transport errors from Claude itself belong to the separate handling path in API errors. Tool errors occur after a model call successfully requests an operation. Keeping those layers distinct allows accurate retry behavior: retrying the Claude request is not the same as retrying the downstream tool.

When a tool throws unexpectedly, record the full diagnostic internally while returning a bounded message to the model. This prevents stack traces and credentials from becoming conversational context while preserving enough evidence for operators.

Scale toolsets without flooding context

Large agents can accumulate dozens or hundreds of integrations. Loading every tool schema into every request consumes context and can make selection harder. Anthropic’s platform now includes approaches such as tool search and programmatic tool calling for larger tool environments, while ordinary applications can also expose only the tools relevant to the current workflow.

Organize tools by capability and trust level. A support workflow might expose customer lookup and ticket update, while a deployment workflow gets repository, CI, and infrastructure tools. Dynamic exposure should still be deterministic from the application’s state; do not let untrusted content decide which privileged toolset becomes available.

The durable pattern is simple: tools are typed capabilities with owners, permissions, limits, and failure behavior. When Claude chooses among those capabilities, the application remains responsible for execution. That separation lets tool calling stay flexible without surrendering the controls that make production systems dependable.

Use polling or callbacks for long-running operations

Some tools start work that cannot finish inside one model turn: deployments, data exports, scans, or batch jobs. A strong pattern is to return a job identifier immediately, then expose a separate status tool. Claude can decide when to check again, while your worker system owns the long-running process and its retries.

If the application can push completion events, resume the workflow from a durable checkpoint rather than holding an HTTP request open. This separates model latency from job latency and prevents a disconnected client from leaving an ambiguous operation behind.

Make job status states explicit—queued, running, succeeded, failed, cancelled—and include bounded result metadata. The model should not infer completion from a vague message or scrape logs to decide whether an operation finished.

Treat tool selection as something you can evaluate

Build a small evaluation set of user requests and expected tool behavior. Some prompts should call a specific tool, some should choose between tools, and some should answer without calling anything. Measure wrong-tool calls, unnecessary calls, invalid arguments, and failure recovery after error results.

When performance is poor, improve the interface before adding more prompt text. Sharper descriptions, distinct names, simpler schemas, examples, or fewer simultaneously exposed tools can make selection easier. Tool quality is partly an API-design problem presented to the model.

Evaluation also helps during platform or model upgrades. Re-run the same cases and compare behavior before deploying. An agent that still produces good prose but starts choosing a broader write tool instead of a narrow read tool has changed in a way ordinary answer-quality tests may miss.

Tool loops also need termination rules. Cap the number of tool rounds, detect repeated identical calls, and stop when the workflow is no longer making progress. A model can otherwise spend time retrying a permanent downstream failure or cycling between two incomplete sources. Return a clear partial result or escalation state when the loop reaches its budget so the application fails predictably instead of consuming unbounded latency and cost.

Keep the model-facing tool name stable when implementation details change. A search capability can move from one backend to another without changing the schema if its semantics remain the same. Conversely, if the meaning or permission model changes, version the tool or create a new one rather than silently redefining an existing contract. Stable semantics make agent behavior easier to evaluate across releases.

Related Posts

• CompTIA Security Operations

• Microsoft AI-103: Choosing Embeddings on Azure

• Microsoft AI-103: Tool Calling in Azure AI Agents

• Microsoft AB-100: Researcher and Analyst in Microsoft 365

• Microsoft SC-500: Passkeys in Microsoft Entra ID

• Amazon AWS AIP-C01: Vector Search for Bedrock RAG

• Anthropic CCAO-F: Production Incident Playbooks for Claude

• Microsoft AZ-104: FSLogix for Azure Virtual Desktop

• CompTIA SY0-701: Identity and Access Control

• Cisco 200-301: Network Automation with RESTCONF