Amazon AWS AIP-C01: Multi-Agent Workflows on AWS
Multi-agent workflows divide a complex task among specialized agents instead of asking one large agent to understand every domain, tool, and permission boundary. On AWS, the current architectural center for new agent development is Amazon Bedrock AgentCore. AWS moved Amazon Bedrock Agents into maintenance mode as “Bedrock Agents Classic” on July 30, 2026, and recommends AgentCore for new agent workloads and future migration.
That platform shift matters for multi-agent design. Bedrock Agents Classic still supports supervisor and collaborator agents for existing customers, but AWS’s current maintenance-mode guidance says advanced multi-agent collaboration on AgentCore should use code-defined orchestration. AgentCore Runtime can host framework-based or custom agents, supports MCP and A2A communication, and provides identity, memory, gateway, observability, and isolated runtime capabilities.
Multi-agent orchestration is therefore a current production pattern inside Generative AI on AWS, but the implementation choice must reflect the present AgentCore direction.
Split agents by real responsibility
Each specialist should own a coherent job, data set, tool set, or permission boundary.
Multi-agent orchestration works best when collaborators are meaningfully different rather than several copies of the same general-purpose agent.
If every agent shares the same tools, identity, and instructions, one well-designed agent may be simpler and faster.
Use AgentCore for new designs
AgentCore Runtime is framework agnostic and can host LangGraph, Strands, CrewAI, custom code, and other agent implementations.
The Runtime also supports MCP and Agent-to-Agent communication, which gives teams several ways to expose specialist agents and tools to a coordinating workflow.
For new accounts, do not base the architecture on Bedrock Agents Classic creation because AWS no longer opens that service to new customers.
Choose supervisor or router behavior
A supervisor can gather outputs from several specialists and synthesize a final answer, while a router can select one specialist and let that agent respond directly.
Routing usually reduces orchestration overhead and latency when the user request belongs clearly to one domain.
Agent coordination should justify any supervisor synthesis with a real need for cross-domain reasoning.
Use agents as tools
A practical AgentCore pattern is to expose a specialist agent behind a tool or MCP interface and let another orchestrator call it when needed.
This creates a clear input/output contract and allows the specialist to be tested and deployed independently.
Keep the contract narrow enough that the caller knows what work is delegated and what result comes back.
Keep identities separate
Different agents can have different permissions even when they collaborate on one user request.
GenAI IAM should preserve which agent may read which data or call which tool rather than giving the supervisor the union of every collaborator’s authority.
AgentCore Identity can help manage workload identities and outbound credentials for agent applications.
Share context deliberately
Not every collaborator needs the entire conversation, private user history, or another agent’s intermediate reasoning.
Pass the minimum task context and stable business identifiers required for the specialist to do its job.
AgentCore Memory can support short- and long-term context, but memory sharing should be designed according to privacy and authorization rather than enabled by convenience.
Parallelize independent specialists
Research, pricing, policy lookup, and technical validation can sometimes run in parallel.
Latency tuning should compare parallel fan-out with the extra model and tool calls it creates.
Parallel agents are valuable when their work is genuinely independent; otherwise they create synchronization and conflict-resolution overhead.
Define conflict and failure handling
Specialists can disagree, time out, return incomplete evidence, or fail authorization.
The orchestrator should know whether to retry, choose another agent, escalate to a human, or return a partial answer with explicit uncertainty.
Do not ask another model to “resolve” a hard policy conflict when the business already has a deterministic rule.
Evaluate the workflow, not each agent alone
A specialist can score well in isolation while the supervisor routes requests incorrectly or combines outputs badly.
Bedrock evaluation and application-level testing should cover routing accuracy, specialist quality, handoff context, latency, cost, and final business outcome.
For AIP-C01 workloads, the durable multi-agent pattern is specialized roles, code-defined orchestration on AgentCore, narrow identities, explicit context sharing, measured parallelism, and end-to-end evaluation. More agents should create clearer responsibility, not more mystery.
Teams migrating from Bedrock Agents Classic should distinguish a feature migration from an architecture rewrite. AWS states that existing Classic agents continue to work, but new features will not be added and new accounts cannot create them. A simple model-plus-tools agent may map cleanly to AgentCore harness patterns, while complex supervisor collaborations may require custom orchestration code in AgentCore Runtime.
The migration is also an opportunity to revisit whether multiple agents are still necessary. Older designs sometimes created collaborators because the managed service made the pattern easy. If one modern model and a smaller tool set can satisfy the same workload with lower latency and fewer handoffs, simplification can reduce both cost and security surface.
Agent-as-tool patterns create useful modularity. A specialist can expose one bounded capability through MCP while retaining its own model, memory, and IAM. The supervisor does not need to know the specialist’s internal implementation; it only needs a clear contract and an authorization path to call it.
A2A communication can support richer agent collaboration, but protocols do not solve trust. Authenticate every agent, identify the caller, constrain which peers may invoke which capability, and validate messages. Treat agent messages as untrusted input even when they originate from another component inside the same AWS account.
AgentCore Runtime session isolation can reduce cross-user state leakage by running sessions in isolated environments. Persistent or shared memory should still be designed explicitly because an authorized memory store can intentionally outlive one runtime session. Keep user, tenant, and agent boundaries visible in the memory model.
Multi-agent observability should show the call graph. Operators need to know which supervisor routed the request, which collaborators ran, how many model calls occurred, which tools were invoked, and which agent produced the final claim or action. Without that lineage, debugging becomes an argument between several autonomous components.
Budget each collaborator. A research specialist might be allowed several retrieval turns while a payment specialist should perform one validated action and stop. Separate budgets make abnormal behavior easier to detect and prevent one collaborator loop from consuming the whole workflow’s latency or spend.
Multi-agent architecture is mature when specialization simplifies ownership. Each agent has a clear job, identity, tool boundary, evaluation set, and failure behavior; the supervisor has a clear routing policy; and the complete workflow is tested as one product. If adding agents makes those answers harder, the design is probably moving in the wrong direction.
AgentCore Instances add another option for persistent or multi-day workloads and can host multiple collaborating agents on shared managed EC2 infrastructure. That can fit long-running research or stateful coordination, while serverless microVM sessions are attractive for isolated request-oriented workloads. Choose the runtime according to session duration, isolation, cost, and collaboration needs rather than assuming one compute model fits every agent.
Shared memory should have an explicit schema and ownership. One collaborator may write a user preference while another writes a business decision or tool result. Without boundaries, long-term memory becomes an unreviewed shared database that can leak context or accumulate contradictory state.
Supervisor routing should be evaluated with confusion cases where two specialists appear plausible. Measure whether the orchestrator asks a clarifying question, routes consistently, or calls both unnecessarily. Routing accuracy can matter more to user experience than the quality of either specialist in isolation.
For existing Bedrock Agents Classic customers, migration can be staged. Keep the current production agent operating while a code-defined AgentCore equivalent is evaluated against the same conversations, tools, and business outcomes. This reduces migration pressure and turns the platform transition into an evidence-backed release rather than a forced rewrite.
Tool access should remain local to the specialist that owns the task. A supervisor does not need direct payment permissions simply because one collaborator can process payments. Keeping tool authority with the specialist makes the collaboration graph easier to audit and limits the impact of supervisor prompt injection.
Use deterministic workflow when the handoff sequence is known. If the process always validates identity, looks up an order, checks eligibility, and creates a return, a Step Functions or application workflow may be clearer than several agents negotiating the sequence. Multi-agent reasoning earns its complexity where task decomposition or specialist interpretation genuinely changes at runtime.
Multi-agent costs should be reported as one business task and by collaborator. This lets teams see whether a specialist is delivering unique value or merely adding another model call. Optimization can then remove unnecessary collaboration without guessing which component created the extra spend.