Microsoft AI-103: Durable AI Workflows with Queues
AI workflows become operationally difficult when one user request expands into minutes or hours of work. A model may call tools, wait for external systems, request human approval, hit a rate limit, or resume after infrastructure failure. If all of that logic lives inside one short-lived request handler, a restart can lose progress and force expensive work to begin again.
Azure now provides two complementary building blocks for this problem. Durable Task gives long-running AI and agent workflows checkpointing, state persistence, automatic recovery, and distributed coordination. Messaging services such as Azure Service Bus provide durable queues and topics that decouple producers from workers, absorb bursts, support retry and dead-letter patterns, and make asynchronous work easier to control.
The key is to use queues and durable orchestration for different responsibilities inside Azure AI engineering.
Use a queue when work should survive the caller
A queue creates a durable handoff. The web request, event handler, or agent can enqueue a work item and return without keeping the original process alive. A worker can then process the message when capacity is available.
This is useful for document processing, evaluation jobs, bulk enrichment, long-running tool calls, notification work, or AI tasks that do not need an immediate synchronous response. Queue depth also becomes an operational signal showing whether demand is outpacing processing capacity.
The broader architecture principle in event-driven systems is loose coupling: producers should not need to know which worker instance will eventually complete the work.
Use durable orchestration when the workflow has state
A queue can hold a message, but it does not by itself remember that a workflow completed step three, is waiting for approval, and should resume at step four after a restart. Durable Task is designed for that stateful coordination.
Microsoft’s current Durable Task guidance for AI agents emphasizes long-running sessions, expensive token consumption, checkpointing, and automatic recovery. The runtime persists workflow state so completed work does not need to be repeated simply because a worker moved or failed.
This fits naturally with agent workflows. The durable orchestrator owns progress; agents and functions perform bounded pieces of work.
Queues should carry commands or events with explicit contracts
A queue message needs enough information for a worker to act without depending on transient in-memory state. Use stable identifiers, operation type, version, correlation ID, authorization context reference, and the location of any large payload rather than stuffing an entire conversation into the message.
Keep message schemas versioned. A worker deployed tomorrow may receive a message created by yesterday’s producer. Backward compatibility matters when queues contain work during deployments.
Do not place secrets or unnecessary sensitive content in messages. A queue is durable by design, so its retention and dead-letter behavior must be part of the data-governance model.
Idempotency protects the workflow from duplicate work
Distributed systems cannot assume that a message or command is processed exactly once at the business level. Network uncertainty and retries can create duplicate delivery or duplicate send attempts. Azure Service Bus supports duplicate detection for configured queues and topics, but application-level idempotency is still valuable.
Assign a stable operation ID and record the business effect associated with it. If a worker receives the same operation again, it should recognize that the action already completed or safely repeat it without creating a second side effect.
This is especially important when the expensive step is an LLM call or external transaction. A duplicate can double token cost or repeat an irreversible business action.
Retry policies should distinguish transient and permanent failure
Rate limiting, temporary service unavailability, or a network timeout may justify retry. Invalid input, denied authorization, or a policy violation usually does not. Retrying permanent failures wastes capacity and can hide a defect behind repeated queue delivery.
Use bounded retries with backoff. After the retry policy is exhausted, move the message to a dead-letter path with enough context for investigation. Operators should be able to see why a message failed and whether replay is safe.
The rate-limit behavior discussed in AI rate limits should be handled as part of this retry design rather than as an exception discovered during load.
Durable checkpoints prevent token waste after failure
An agent workflow may already have performed retrieval, model reasoning, and tool calls before a later dependency fails. Without checkpointing, recovery can repeat all of that work. Durable execution records progress so the workflow can resume from a known state.
This matters economically as well as operationally. Token consumption is not recovered when a process crashes. Checkpointing protects work that has already been paid for.
State should remain intentional. Persist the minimum workflow data required for reliable continuation rather than treating durability as permission to retain every prompt and tool payload forever.
Human approval becomes an external event
Long-running AI workflows often need to pause for a person. Durable orchestration can wait without consuming an active worker, then resume when an approval or rejection arrives. The approval event should carry the identity of the approver, decision, timestamp, and operation being approved.
The model can prepare a recommendation, but the workflow should enforce the gate. That prevents an agent from bypassing approval simply because its prompt changed.
This pattern is particularly useful for high-impact tool calls in multi-agent workflows, where several agents may contribute before one controlled action is permitted.
Backpressure keeps AI capacity from becoming the bottleneck
A queue absorbs bursts and lets workers process at a rate the downstream model or tool can sustain. This is a form of backpressure. Instead of every producer immediately calling the model and competing for quota, work accumulates visibly.
Autoscale workers from queue depth only when model quota and downstream capacity can support the additional concurrency. Scaling compute without scaling AI capacity can increase throttling rather than throughput.
Capacity planning should therefore include queue arrival rate, processing time, retry rate, and acceptable backlog age.
Choose the smallest durable mechanism that fits the workflow
Not every asynchronous task needs a full orchestrator. A single queue and stateless worker can be enough for one-step work. Durable Task becomes valuable when the process has multiple steps, long waits, human interaction, fan-out/fan-in, or recovery requirements that would otherwise be reimplemented manually.
Microsoft’s current guidance also distinguishes Durable Task from agent frameworks: Durable Task provides reliable execution and can work with Microsoft Agent Framework or other agent stacks. The agent framework handles agent behavior; the durable runtime handles persistence and recovery.
For engineers working through AI-103, that separation is a durable architecture skill. Queues decouple work. Durable orchestration remembers progress. Together they let AI workflows survive time, failure, and fluctuating capacity without turning every process into one fragile synchronous request.
Message settlement is part of reliability. A worker should not permanently remove a queue item before the business step is safely complete. Lock-based processing lets a worker receive work, complete it after success, or abandon it for retry. Long AI operations need enough lock time or renewal so a slow model call is not mistaken for a dead worker.
Ordering should be added only where the business process requires it. Global ordering can reduce throughput. Service Bus sessions can preserve order for related messages, such as operations for one case or conversation, while independent sessions continue in parallel. This is usually more scalable than forcing the entire queue through one sequence.
Dead-letter queues need an operating process, not only a setting. Record the failure reason, workflow version, correlation ID, and retry count. Decide which messages can be corrected and replayed and which represent permanent rejection. Durable messaging creates operational value only when failed work has an owner and a safe resolution path.
Observability should connect queue state to workflow state. Queue age can show that work is waiting, while durable-orchestration telemetry shows where active workflows are paused or failing. Looking at only one layer can produce the wrong diagnosis: a shallow queue can hide many stuck workflows, and a large queue can be healthy if workers are deliberately rate-limited to protect downstream quota.
Define service objectives around the business outcome, such as time from event receipt to completed AI task, rather than only message-processing speed. This keeps reliability work focused on what users and downstream systems actually experience.
Poison-message handling should be explicit. If the same payload repeatedly fails because of bad data or an incompatible schema, retries only waste capacity. After a bounded number of attempts, isolate the work, preserve the diagnostic context, and require a deliberate replay decision rather than letting the system loop forever.