Microsoft AI-103: Choosing Azure AI Deployment Models
Azure AI deployment choices determine more than where a model runs. They affect where inference data can be processed, how capacity is allocated, whether billing is pay-per-token or reserved, how much latency variation to expect, and whether the workload is suited to real-time or asynchronous processing. Choosing the wrong deployment type can create a compliance or reliability problem even when the model itself is a good fit.
Microsoft Foundry currently uses serverless API as the preferred deployment option for a broad set of Foundry Models. Within that option, deployments can be global, data-zone, geography-based, provisioned, batch, or developer-oriented depending on the model. Foundry also supports managed compute for open-source, partner, and custom models that require dedicated GPU-backed hosting.
The decision is part of the current AI-103 skill set because Azure AI engineers are expected to choose deployment options, configure model deployments, integrate CI/CD, and operate solutions. Deployment is architecture, not packaging.
Start with the data-processing boundary
The first question is where inference processing is allowed to occur. Global deployments can route across Azure regions to provide broad capacity and early model availability. Data Zone deployments keep processing within a Microsoft-defined zone such as the United States, European Union, or Asia Pacific. Geography-based Standard or Regional Provisioned deployments keep processing within the selected Azure geography.
Data at rest and inference processing are different considerations. A resource can live in one geography while a globally routed deployment processes prompts and responses elsewhere. Architecture review should therefore examine the deployment type itself, not infer behavior from the resource’s home region.
If policy requires a specific processing boundary, apply that constraint before comparing price or throughput. Compliance is a hard filter, not a preference.
Global Standard is the practical default for many workloads
Microsoft recommends Global Standard as a starting point for many general workloads because it offers broad availability, access to new models early, and high default quota. Billing is pay-per-token, which fits variable demand without requiring reserved capacity.
The tradeoff is global processing and best-effort performance. High sustained volume can experience more latency variation than provisioned throughput. A workload with strict geography requirements or highly predictable latency needs may therefore need a different deployment type.
Use Global Standard when its data-processing model is acceptable and the application benefits from elasticity and broad model availability. Do not use it automatically when governance requires tighter locality.
Data Zone Standard balances locality and shared capacity
Data Zone Standard keeps processing within a broader zone rather than a single geography while preserving pay-per-token operation. It can be a useful middle ground for organizations that cannot use worldwide processing but do not require a single-region boundary.
The available zones and supported models change over time, so deployment planning should verify current availability. The zone must also align with the organization’s interpretation of data residency and regulatory requirements; “EU zone” and “single EU region” are not the same technical promise.
Where a zone-level boundary is acceptable, Data Zone Standard can offer more routing flexibility than a geography-constrained Standard deployment.
Standard is for geography-constrained pay-per-token workloads
Standard deployments keep processing within the Azure geography and use pay-per-token billing. This can fit lower-to-medium volume applications that need geography-level data processing but do not require reserved throughput.
The cost of tighter geography is usually a smaller capacity pool and potentially later availability for newly released models. Teams should verify model support and quota before promising an architecture around a specific geography.
This is a common place where Azure network access and data-processing choices become confused. Private networking controls how the application reaches the service; deployment type controls where inference may be processed. Both can matter, but they solve different problems.
Provisioned throughput is for predictable performance
Provisioned deployment types reserve model processing capacity through provisioned throughput units. Global Provisioned uses global routing, Data Zone Provisioned stays within a zone, and Regional Provisioned stays within the Azure geography. The common benefit is dedicated capacity with more predictable throughput and lower latency variance.
Provisioned throughput is most attractive when demand is sustained enough to justify reservation and the service objective values consistency. It is less attractive for uncertain prototypes or highly bursty workloads that may leave reserved capacity unused.
Capacity has to be sized for the specific model because PTU requirements and throughput vary. Capacity planning should therefore happen before a provisioned purchase, not after.
Batch deployments trade responsiveness for economics
Global Batch and Data Zone Batch are designed for large asynchronous jobs rather than interactive requests. They can offer substantial cost savings in exchange for delayed completion. This is appropriate for offline evaluation, content processing, summarization, enrichment, or other work that does not need a synchronous response.
Batch should be treated as a separate workload class. Do not put interactive requests into a delayed queue simply because the unit price is lower. Instead, identify work that is naturally asynchronous and move it out of the real-time capacity pool.
Batch can also simplify peak planning by shifting nonurgent jobs away from periods when user-facing traffic is high.
Developer deployments are for evaluation, not production
Developer deployment types are intended for evaluating certain fine-tuned models. They have important limitations such as short lifetime, no availability SLA, and no data-residency guarantee. Those properties make them useful for experimentation and poor as a production serving tier.
This distinction should appear in deployment policy so an experiment does not become a hidden dependency. Production promotion should require moving the model to a deployment type with the necessary SLA, residency, and capacity characteristics.
Temporary evaluation infrastructure is valuable precisely because it is cheap and easy to discard. Treat it as temporary.
Managed compute solves a different hosting problem
Serverless API deployment covers a wide range of Foundry Models, but not every model or customization fits that path. Managed compute provides dedicated GPU-backed hosting for open-source, partner, or custom models where the platform needs to manage compute on the team’s behalf.
Managed compute introduces classic infrastructure considerations such as instance sizing, scaling, image and environment management, and capacity availability. It can offer flexibility that serverless does not, but the operational burden is higher.
The choice between serverless and managed compute should therefore be driven by model compatibility and control requirements, not by a belief that dedicated compute is inherently more production-ready.
Deployment type and model selection must be evaluated together
A model may look ideal until the team checks where it can be deployed. Some models support only certain regions or SKUs. New models often appear first in broader global deployment types and reach geography-constrained options later. The deployment requirement can therefore eliminate model candidates.
That is why model selection should treat deployment compatibility as a hard constraint. Evaluate only the models the organization can actually run under its data, latency, support, and capacity requirements.
The reverse is also true: a deployment strategy built around one model family can become fragile when that model approaches retirement. Maintain an evaluation process and migration path rather than binding the application permanently to one version.
A short decision tree keeps the deployment choice practical.
If there is no special processing-location requirement and demand is variable, Global Standard is a strong starting point. If processing must stay in a defined zone, consider Data Zone Standard. If it must remain within an Azure geography, evaluate Standard. If demand is sustained and predictable with strict latency needs, move the same boundary decision into the appropriate provisioned type. If the job is asynchronous, evaluate Batch. If the model requires dedicated hosting, evaluate managed compute.
Then validate current model availability, quota, pricing, SLA, and regional support. Those details change faster than the decision logic.
For teams building in Azure AI engineering and working through Azure AI developer certification, the durable skill is not memorizing SKU names. It is understanding the four questions every deployment answers: where data is processed, how capacity is supplied, how the workload is billed, and what operational guarantees the application needs.
High availability should be reviewed separately from data-processing scope. A globally routed deployment can draw on broad infrastructure, but an application can still be affected by a resource, networking, identity, or dependency failure in its own architecture. Conversely, a geography-constrained deployment can meet residency policy while requiring the team to design its own cross-region recovery strategy. Deployment type is one layer of resilience, not the entire resilience plan.