Practice Exams:

Microsoft AZ-104: ExpressRoute Resiliency Patterns

ExpressRoute is built with redundant connectivity inside Microsoft’s network, but one circuit at one peering location does not protect the entire hybrid path from every failure. Enterprise resiliency depends on circuit design, provider diversity, peering-location diversity, on-premises edge devices, cross-connects, BGP sessions, virtual network gateways, regional topology, and fallback paths. The business outcome is private connectivity that continues through maintenance, device failure, provider failure, and—where required—peering-site or regional failure.

Microsoft’s current ExpressRoute guidance describes Standard, High, and Maximum resiliency options. Current Well-Architected guidance recommends Maximum resiliency for critical workloads: multiple circuits in different physical peering locations, diverse provider/network paths, zone-redundant ExpressRoute gateways, and tested failover. Microsoft also documents ExpressRoute Metro as a high-resiliency option in supported metropolitan locations and a public-preview Resiliency Guard for guided multi-homed gateway configuration.

ExpressRoute resiliency is therefore a core hybrid-network topic inside Azure Architecture in Practice.

Understand standard circuit redundancy

Every ExpressRoute circuit includes redundant connections to Microsoft Enterprise Edge routers at one peering location.

This protects against individual link or device failure inside that location.

ExpressRoute versus VPN should still recognize that one circuit remains dependent on the chosen provider path and peering site.

Use multiple peering locations for maximum resiliency

For critical workloads, Microsoft recommends circuits at different physical ExpressRoute locations so a site-wide failure does not remove all private connectivity.

Terminate those circuits through diverse on-premises equipment and, where practical, different provider networks.

Two circuits that share the same carrier building, router, or fiber path can provide less real diversity than the Azure configuration suggests.

Use ExpressRoute Metro where it fits

ExpressRoute Metro provides location-level resiliency across two peering locations within supported metropolitan areas through a single Metro construct.

It can simplify high-resiliency architecture where the service is available and matches on-premises topology.

Architects should compare Metro with geographically separate circuits when site-failure and regional-failure requirements extend beyond one metro area.

Deploy zone-redundant gateways

ExpressRoute virtual network gateways connect circuits to Azure virtual networks.

Current Microsoft reliability guidance recommends zone-redundant gateway SKUs for production so gateway instances span Availability Zones.

Zone resilience protects the Azure gateway layer but does not replace peering-location and provider diversity outside the region.

Use active-active paths

Design both ExpressRoute circuit connections to carry production routes rather than keeping one path completely untested until failure.

Active-active routing validates the secondary path continuously and can distribute traffic where the network supports it.

BGP attributes, routing policy, and on-premises architecture should make failover predictable instead of relying on emergency manual changes.

Consider VPN as a backup path

Microsoft documents site-to-site VPN as a possible backup for ExpressRoute private peering when business requirements accept its internet path, bandwidth, and latency characteristics.

The backup should be preconfigured and tested; an incident is a poor time to design encryption, routing, and firewall policy from scratch.

Use routing preferences that keep normal traffic on ExpressRoute and fail over deliberately when the private path is unavailable.

Design BGP for fast failure detection

BGP session behavior and technologies such as Bidirectional Forwarding Detection can influence how quickly failures are detected and routes withdrawn where supported.

Fast control-plane convergence is useful only if applications tolerate the brief traffic interruption and the alternate path has enough capacity.

Measure failover at the application level instead of assuming a BGP timer proves the user experience.

Plan gateway and route scale

ExpressRoute gateways have route, flow, throughput, and VM-scale limits that vary by SKU.

Large enterprises should summarize routes and select a gateway whose limits cover normal and failure-state traffic.

Azure Route Server can add dynamic NVA integration, but it also contributes route propagation and scale considerations to the hybrid design.

Test realistic failures

Disconnect one circuit, one provider path, one on-premises router, and one gateway dependency in controlled tests.

For AZ-700, the durable resiliency model is circuit redundancy → site diversity → provider diversity → zone-redundant gateway → active routing → fallback path → measured application recovery.

Hybrid connectivity is resilient only when every layer in the path has been considered and the alternate path has actually carried production-shaped traffic.

Maintenance readiness should be part of design. Microsoft and connectivity providers perform planned work, and customer-controlled maintenance options exist for supported gateways. A resilient architecture should be able to lose one path for maintenance without requiring an outage window for every hybrid application.

Route advertisement should remain symmetric and consistent across redundant circuits unless traffic engineering intentionally prefers one path. Missing prefixes on the secondary path can remain hidden for months until the primary fails. Automated route comparison can detect this drift early.

Regional recovery should use ExpressRoute topology that can reach the secondary Azure region. A redundant circuit into one region does not automatically provide a complete disaster-recovery path if on-premises routing, gateway connections, or application DNS cannot reach the recovery environment.

Operational dashboards should show circuit state, BGP sessions, gateway health, route counts, provider incidents, and effective application reachability. The objective is to detect degraded redundancy before the second failure turns it into an outage.

The strongest ExpressRoute design has enough diversity that no single customer router, provider path, peering location, Azure gateway zone, or untested backup assumption can remove private connectivity for the workloads whose business target requires it.

ExpressRoute circuit redundancy begins at the provider edge. A single circuit includes redundant connections to Microsoft routers at one peering location, but the customer side can still contain one router, one cross-connect, one carrier, or one building. Map the complete physical path so “two links” does not hide a shared failure domain before Microsoft network.

Maximum resiliency should include peering-location diversity and customer-edge diversity. Microsoft’s current guidance recommends multiple circuits in distinct locations for critical workloads. Terminate those circuits on independent routers and, where possible, separate physical facilities or provider networks so one fiber cut or provider maintenance cannot remove both paths.

High Resiliency through ExpressRoute Metro can be attractive in supported locations because it spans two peering sites in the same metro. The architecture should still ask whether the business needs protection only from one peering-site failure or from broader metro/regional disasters. Maximum resiliency across distinct locations can provide a larger failure boundary.

ExpressRoute gateways must be sized for failure-state load. If two circuits normally share traffic, the remaining path and gateway capacity must handle the full workload after one path fails. Review throughput, packet rate, flow count, learned routes, and gateway SKU limits in the degraded state rather than only during normal operation.

Zone-redundant ExpressRoute gateways reduce the risk of an Azure Availability Zone outage affecting the gateway layer. Use current AZ-enabled SKUs for production where supported and avoid legacy non-zone-redundant gateway designs when the workload’s availability target requires zone protection.

Routing policy should prevent accidental active/standby drift. If the secondary circuit has less-preferred routes, test those prefixes regularly and monitor that the advertisements remain complete. Configuration drift on a backup path can remain invisible for months because no production traffic exercises it.

BFD can improve failure detection on supported ExpressRoute paths, but fast detection is only useful if the alternate path converges and has enough capacity. Tune failure detection with provider and equipment guidance and test application behavior during the resulting route transitions.

Site-to-site VPN backup should have realistic expectations. Internet VPN can preserve critical management or business traffic when ExpressRoute fails, but it may not match the bandwidth or latency of the private circuit. Prioritize essential prefixes or applications if the backup path cannot carry normal peak demand.

ExpressRoute Direct customers should include redundant physical ports and cross-connects in the same resiliency analysis. Direct reduces provider abstraction but increases the customer’s responsibility for physical diversity, MACsec or encryption options, router configuration, and maintenance readiness.

Route filtering should be conservative. Over-advertising on-premises prefixes increases route scale and can create unintended transit. Summarize where safe, advertise only what Azure needs, and verify that secondary circuits and recovery regions learn the same critical routes.

Monitoring should include synthetic traffic across both circuits, not only BGP session state. A session can remain established while a provider experiences performance degradation or one application prefix is missing. End-to-end probes to critical Azure endpoints can reveal partial failures earlier.

Change management should schedule circuit, provider, gateway, or router maintenance so redundancy remains intact. If one path is already degraded, postpone planned work on the remaining path unless the business explicitly accepts the risk. Resiliency is an operating discipline as much as a topology.

Regional disaster recovery should be connected to circuit architecture. The secondary Azure region needs a reachable ExpressRoute gateway, routes, and DNS/application failover. Test the actual DR traffic path from on-premises users and systems, not just Azure-to-Azure recovery.

For enterprise architecture, the strongest ExpressRoute pattern has no hidden single point of failure from the workload to Microsoft: diverse customer routers, provider paths, peering locations, circuits, zone-redundant gateways, route advertisements, and a tested fallback strategy appropriate to the service objective.

Keep peering configuration identical enough across redundant circuits that failover does not change application reachability unexpectedly. Route filters, communities, private-peering prefixes, and customer-edge policy should be compared regularly so the secondary path is genuinely equivalent for the traffic it is expected to carry.

Related Posts

• CompTIA Security Operations

• IT Operations & Project Delivery

• Microsoft AI-103: Building Multi-Agent Workflows on Azure

• Microsoft AI-103: From AI Prototype to Production on Azure

• Microsoft AI-103: Serverless Patterns for Azure AI

• Microsoft AB-100: Integrating Agents with Power Platform

• Microsoft SC-500: KQL for Security Investigations

• Amazon AWS AIP-C01: Secrets Management for GenAI Apps

• Anthropic CCAO-F: Claude Governance for Regulated Teams

• Microsoft AZ-104: Cost Governance for Azure Subscriptions