Practice Exams:

Microsoft AI-103: Capacity Planning for Azure AI

Capacity planning for Azure AI begins with a simple correction: serverless does not mean unlimited. Model APIs can have tokens-per-minute limits, requests-per-minute limits, concurrency constraints, regional availability, deployment-specific quotas, and capacity that changes by model and SKU. A workload that looks small in request counts can still be large in tokens, while a high-request workload with tiny prompts can hit request limits before token limits.

Microsoft Foundry separates standard pay-per-token deployment types from provisioned throughput and batch options. Azure OpenAI quotas are also scoped by factors such as subscription, region, model, and deployment type. Planning therefore has to translate application behavior into the quota dimension the platform actually enforces.

The current AI-103 responsibilities include deployment, optimization, monitoring, and scaling. Capacity is the bridge between those topics. It belongs inside Azure AI engineering from the design stage, not after the first production throttling event.

Measure tokens and requests separately

Requests per minute and tokens per minute describe different workload shapes. A chat application might receive many short requests and hit RPM first. A document-analysis workflow might submit fewer requests with long prompts and large completions, consuming TPM much faster.

Estimate input and output tokens separately. Input volume is affected by system prompts, conversation history, retrieved documents, tool descriptions, and user content. Output volume depends on the task and generation limits. Long context can dominate throughput even when the visible user prompt is small.

Use percentiles rather than only averages. The average request may contain 2,000 tokens while the 95th percentile contains 20,000. Capacity built around the average can fail precisely on the difficult requests users care about most.

Convert user traffic into a workload envelope

Start with business demand: active users, requests per user, peak concurrency, session length, scheduled jobs, and geographic distribution. Convert that demand into model requests, then tokens. A single user action may generate multiple model calls if the workflow uses routing, retrieval reformulation, agent planning, validation, or retries.

Write down three operating points: normal, peak, and degraded. Normal describes common load. Peak covers known surges. Degraded describes what the application should still accomplish when quota is constrained or a dependency is slow.

This follows a basic serverless capacity principle. Elastic infrastructure removes some provisioning work, but it does not remove limits, budgets, or the need for demand modeling.

Quota is not the same as guaranteed throughput

Standard deployments expose quota and rate limits but generally provide best-effort service. Having TPM quota does not mean every high-volume request will have identical latency. Shared infrastructure, request shape, and model behavior can still produce variability.

Provisioned throughput is designed for workloads that need predictable capacity and lower latency variance. Teams purchase or reserve provisioned throughput units appropriate to a model and deployment type. That creates a different planning problem: instead of hoping shared capacity absorbs a sustained load, the team sizes and pays for dedicated capacity.

The choice should be based on workload consistency and service objectives. Bursty or uncertain demand can fit standard deployments well. Sustained high-volume or latency-sensitive applications may justify provisioned throughput.

Deployment geography changes the capacity options

Global Standard can dynamically route across Azure infrastructure and generally offers broad model availability and high default quota. Data Zone options constrain processing to a defined zone such as the United States, European Union, or Asia Pacific. Geography-based Standard keeps processing within the Azure geography. Provisioned variants add reserved capacity to some of those scopes.

Data-processing policy can therefore reduce the capacity pool available to a workload. A team may prefer a global option for elasticity but be required to use a geography-constrained deployment. That decision has to be made before the capacity model is finalized.

The article on deployment models compares these deployment types directly. Capacity planning should consume that architecture decision rather than treat all deployments as interchangeable.

Retries can turn throttling into a traffic storm

A naive retry policy makes a capacity problem worse. When requests receive rate-limit or transient errors, every client retries at once, increasing pressure on the service. Good retry behavior uses exponential backoff, jitter, bounded attempts, and awareness of server-provided retry guidance.

Centralized queues or application-level concurrency limits can protect the model endpoint from synchronized bursts. Delay-tolerant work can move to batch processing instead of competing with interactive requests. User-facing applications can degrade gracefully by reducing optional calls, shortening context, or postponing nonessential enrichment.

Capacity engineering is therefore partly demand shaping. The application controls how aggressively it sends work and which requests receive priority when resources are constrained.

Context windows create hidden capacity costs

Long-context models make it possible to send large histories and documents, but every token still has processing cost and may count toward rate limits. Repeatedly sending an entire conversation or document set can consume capacity even when the user asks a short question.

Use retrieval to bring only relevant evidence into the prompt. Summarize or compact long histories when the task allows it. Keep tool descriptions concise. Cache stable system context where the platform supports it. Remove duplicated instructions.

This is another point where Azure AI Search and capacity planning meet. Better retrieval can reduce prompt size while improving evidence quality, which helps both cost and throughput.

Model routing can improve capacity efficiency

Not every request needs the same model. Classification, extraction, routing, or simple transformations may run well on a smaller model, while complex reasoning uses a larger one. Splitting traffic by task can reduce token cost and release capacity on the expensive deployment.

The routing logic must be measured. A cheap model that misroutes requests can create retries or poor outcomes that erase the savings. Track success rate and downstream cost by route.

Model selection should therefore include capacity efficiency. The model selection process is stronger when it evaluates quality per successful task, not only answer quality in isolation.

Plan capacity for safe releases and failure recovery

Steady-state demand is not the maximum capacity requirement. Blue-green deployment can temporarily run two versions. Shadow traffic duplicates requests. Canary rollout requires headroom for both baseline and candidate. A regional failure may shift users or jobs to another deployment.

Include those events in the workload envelope. If the baseline normally consumes eighty percent of available quota, there may be no safe room for a mirrored release or recovery surge. Reliability requires unused capacity.

This is why blue-green releases and canary releases should be reviewed with quota data before the release begins.

Monitor the signals that explain saturation

Dashboards should distinguish request rate, input tokens, output tokens, throttling, latency, concurrency, errors, retry volume, and queue depth. Aggregate utilization alone is not enough. If latency rises while token rate is stable, the problem may not be quota. If retries climb after throttling, the application may be amplifying the incident.

Tag metrics by deployment, model, workload, and environment. Capacity issues are easier to diagnose when an operator can see that one agent workflow or one batch job is consuming the majority of demand.

For teams following Azure AI developer certification, capacity planning is a practical engineering skill: know the workload, know the quota scope, shape demand, reserve headroom, and choose provisioned capacity only when the workload and service objective justify it. The goal is not maximum quota. It is predictable service under the traffic the application actually receives.

Capacity planning should also become an ongoing feedback loop rather than a one-time forecast.

The first capacity estimate will be wrong because production traffic always contains behaviors the design model did not anticipate. Treat the estimate as a baseline and reconcile it with observed data after launch. Compare forecast and actual tokens per request, retry rate, peak concurrency, and user growth. When the difference is material, update the model rather than simply requesting more quota.

Capacity reviews should also follow product changes. Adding RAG increases input tokens. Adding a second agent can multiply calls. Increasing maximum output length changes completion demand. A new safety or evaluation service adds its own latency and quota. Release review should therefore ask whether the change modifies the workload envelope, not only whether it changes application logic.

This feedback loop helps teams distinguish genuine demand growth from avoidable inefficiency. The cheapest capacity is often the work the application no longer sends because context was trimmed, retries were fixed, batch jobs were rescheduled, or a simpler model handled a routine task.

Related Posts

• AWS Architecture in Practice

• AWS Cloud Operations

• CompTIA Security Operations

• Data & AI on Google Cloud

• IT Operations & Project Delivery

• IT Support with CompTIA

• ServiceNow Platform Engineering

• Microsoft AI-103: Agent Identity in Azure AI Foundry

• Microsoft AI-103: Building Multi-Agent Workflows on Azure

• Microsoft AI-103: Canary Releases for AI Models