Practice Exams:

Microsoft AZ-104: Availability Zones and Failure Domains

Azure Availability Zones are separate groups of datacenters within a region, with independent power, cooling, and networking designed to reduce the chance that one local failure affects every zone at once. Zones provide a failure-domain boundary, but they do not automatically make every workload zone resilient. The result depends on whether each Azure service is deployed as zone-redundant, zonal across multiple zones, or nonzonal, and on how the application handles traffic, state, dependencies, and failover.

Microsoft’s current reliability guidance distinguishes zone-redundant resources, where the service distributes or replicates across multiple zones and Microsoft generally manages failover, from zonal resources pinned to one zone, where the customer must deploy and coordinate multiple resources to survive a zone loss. Microsoft recommends zone-redundant deployment where possible for production workloads unless a clear requirement justifies a zonal choice.

Availability-zone design is therefore a core topic inside Azure Architecture in Practice.

Understand the zone boundary

Availability Zones are within one Azure region and are connected with low-latency networking while maintaining independent infrastructure.

Azure regions and zones should be treated as operational geography, not just resource-placement labels.

A zone failure is smaller than a region failure, so zone resilience and regional disaster recovery solve different problems.

Prefer zone-redundant services where possible

Zone-redundant services distribute resources or data across multiple zones and generally manage zone failover within the service.

This can simplify application resilience because the customer does not have to orchestrate individual zonal instances.

Check service-specific reliability guidance because zone redundancy may depend on SKU, region, configuration, or deployment mode.

Use zonal resources deliberately

A zonal resource is pinned to one selected availability zone and is isolated from faults in other zones.

That isolation can be useful for latency, co-location, or service-specific requirements, but it is not automatically resilient to failure of its own zone.

Azure availability options should be chosen according to the failure domain the workload needs to survive.

Build multi-zone architecture around zonal resources

If the service is zonal, deploy equivalent resources in two or more zones and design traffic distribution, data replication, failover, and health checks explicitly.

Virtual machines are a common example: multiple VMs or scale-set instances across zones can provide zone resilience while one VM in one zone cannot.

Test whether stateful components and dependencies fail over as cleanly as stateless compute.

Do not confuse zones with availability sets

Availability sets separate VMs across fault and update domains within the datacenter infrastructure model but do not provide the same physical zone-level isolation.

Availability Zones provide a larger failure boundary within the region.

Choose the mechanism according to the current service architecture rather than assuming older availability-set patterns provide zone resilience.

Align dependent components

A multi-zone application can still fail if its database, firewall, NAT gateway, secrets store, or load balancer remains a single-zone or nonzonal dependency.

Well-Architected design should trace the user flow through every dependency and identify the narrowest failure domain.

One unprotected dependency can reduce the effective resilience of the entire workload.

Balance inter-zone latency and resilience

Some chatty workloads may benefit from keeping tightly coupled components in the same zone for latency, while resilience requires copies in another zone.

Architects should measure the actual inter-zone sensitivity and use scale units or partitioning when appropriate.

Do not sacrifice zone resilience solely because one synthetic benchmark showed a small latency difference without user-impact evidence.

Test zone failure behavior

Failure drills should verify traffic rerouting, data consistency, session impact, autoscaling, alerts, application retries, and recovery after the zone returns.

Managed zone-redundant services reduce operational work, but the complete application still needs to be tested under a zone-loss scenario.

Recovery should be measured from the user’s perspective rather than from one Azure resource’s status.

Keep regional recovery separate

A zone-resilient workload can still be unavailable during a full regional outage.

For workloads whose business targets require regional resilience, pair zone design with a second-region strategy and defined data replication/failover.

For AZ-305, the durable mental model is zone-redundant where possible, multi-zone design for zonal resources, dependency review across the whole flow, then regional recovery where the service-level target requires it.

Capacity is another zone-design concern. A service may be technically zone-redundant but unable to absorb full load after one zone fails if remaining capacity is already heavily utilized. Size scale units and quotas with the expected failure state in mind, not only normal operation.

Deployment processes should also avoid concentrating new resources accidentally. Infrastructure-as-code should specify zone strategy explicitly where the service exposes it, and review should catch workloads whose production resources remain nonzonal in a zone-enabled region without a documented reason.

Monitoring should identify both individual-zone health and workload symptoms. Azure Service Health can provide platform context, while application telemetry confirms whether users are actually affected. One zone incident might be fully masked by redundancy, while a dependency failure elsewhere creates the real outage.

Zone design is mature when teams can explain exactly what survives one zone loss, what degrades, what must fail over manually, and what still requires a regional disaster-recovery plan.

Zone resilience should begin with a component inventory. Classify each resource as zone-redundant, zonal, or nonzonal in the selected region and SKU. Some Azure services support multiple modes, while others have service-specific rules. The architecture should identify any component that cannot meet the workload’s zone-failure target and decide whether to redesign, replicate, or accept the limitation.

Data replication semantics matter as much as compute placement. A service can run application instances in three zones while storing state in one zonal disk or database. Review synchronous versus asynchronous replication, write availability during failure, consistency, and data-loss expectations so the RPO matches the business requirement.

Traffic distribution should also survive zone loss. Load balancers, gateways, DNS, and health probes need to remove failed instances quickly and preserve capacity in the remaining zones. A user flow can remain unavailable after compute failover if the entry point continues sending requests toward a failed zone or if health checks test the wrong dependency.

Quota planning is easy to overlook. If one zone disappears, the remaining zones must have enough compute, networking, database, and service quota to absorb load. Design normal utilization so the system has failure-state headroom, or define automated scale-out that has already been tested under constrained-zone conditions.

Deployment architecture should avoid correlated failure. Spreading application instances across zones but deploying all CI/CD agents, secrets access, or message-processing capacity in one zone can still stop recovery. Review control-plane and operational dependencies as part of the same failure-domain exercise.

Some workloads deliberately use zone-aligned components to reduce latency. When doing so, use multiple independent scale units, each aligned within a zone, and route across units. This pattern can preserve low intra-zone latency while still allowing another scale unit to serve traffic after a zone failure.

Availability Zones do not remove the need for backups. Zone replication protects against infrastructure failure but does not necessarily protect against logical corruption, accidental deletion, ransomware, or bad deployment. Backup and point-in-time recovery solve different failure modes and should be tested independently.

Operational procedures should state whether failover is automatic, application-managed, or manual for each critical resource. Zone-redundant managed services may fail over automatically, while zonal architectures require workload logic or orchestration. During an incident, responders should not have to discover which layer is responsible for moving traffic.

Test failback as well as failover. After the zone recovers, decide whether traffic should rebalance automatically, whether data needs resynchronization, and whether returning capacity creates another disruption. Unplanned failback can turn one outage into two if the architecture assumes recovery is instantaneous.

The key design principle is to align failure domains with service-level targets. Use zones to isolate local infrastructure faults, use region-level recovery for regional disasters, use backups for data loss, and use application retries and observability for transient failures. Each mechanism solves a different class of failure.

Zone placement can also affect maintenance behavior. Microsoft generally stages platform updates to reduce simultaneous impact across zones, but workload teams should still use multiple instances and health-based routing so planned maintenance is absorbed the same way as an unexpected fault.

Stateful services need clear write behavior during partition or failover. Some managed services preserve synchronous consistency across zones; others use replicas with different semantics. Read the service-specific reliability guide and test the application’s retry logic so transient failover does not become duplicate writes or data corruption.

Network appliances and private endpoints should be included in the zone map. A workload can deploy VMs across zones while a single zonal NVA, NAT path, or DNS dependency remains central. Review packet paths as carefully as compute placement.

Document zone strategy in infrastructure-as-code and architecture diagrams so future deployments do not silently regress to single-zone placement. Resilience should be reproducible rather than dependent on one engineer selecting the right portal option.

Review the zone strategy whenever a service tier, region, or dependency changes. A component that was nonzonal at initial launch may later support zone redundancy, creating an opportunity to reduce risk without redesigning the entire workload.

Related Posts

• Azure AI Engineering

• Microsoft Business AI Systems

• Microsoft AI-103: Azure AI Foundry Model Selection

• Microsoft AI-103: Handling Hallucinations in Azure AI

• Microsoft AI-103: Python SDK Patterns for Azure AI

• Microsoft AB-100: Building an AI Champions Program

• Microsoft DP-600: Eventstreams for Real-Time Analytics

• Microsoft SC-500: Threat Modeling Cloud and AI Systems

• CompTIA CS0-003: XDR and SIEM Working Together

• ServiceNow CIS-DF: CI Relationships That Support Operations