Practice Exams:

HA and Disaster Recovery for Enterprise Firewalls

 

High availability and disaster recovery are related, but they solve different failure scopes. A firewall cluster can protect against a device failure without protecting against a building outage, carrier failure, routing mistake, corrupt policy deployment, or regional event. Disaster recovery can provide an alternate site without preserving active sessions or the same public addresses. Treating the two as synonyms creates designs that look redundant while sharing critical dependencies.

This topic originated in the FCSS_EFW_AD-7.6 sequence. Fortinet retired the Enterprise Firewall 7.6 Administrator exam on July 15, 2026, but the current secure-networking path continues to depend on resilience concepts. The technology names may shift; the architecture questions remain: which failures must be survived, how quickly, with what loss of state, and at what operational cost?

A useful design begins by listing failure domains. Appliance hardware is one. Power, switching, routing, internet circuits, data centers, management platforms, certificate services, DNS, identity, and human change are others. HA and DR should be combined to cover the failures that matter to the business rather than duplicated blindly at every layer.

HA protects a service instance; DR protects a business capability

Local firewall HA is usually designed to keep the same security service available when one unit fails or requires maintenance. The cluster shares configuration and may synchronize sessions so the surviving member can continue processing traffic with minimal interruption.

Disaster recovery assumes a larger boundary has failed. The alternate firewall may be in another site or region with different upstream providers, routes, addresses, and application endpoints. Restoring service may require DNS changes, route advertisement, VPN re-establishment, or application failover in addition to firewall activation.

These objectives should have separate recovery targets. Some applications may tolerate reconnecting sessions in seconds, while others can tolerate minutes of network recovery but not hours of data-center unavailability.

Cluster design depends on the surrounding network

Two firewalls do not create high availability if both connect to one switch, one carrier handoff, one power distribution unit, or one routing path. The adjacent infrastructure must provide redundant connectivity and predictable forwarding when the active member changes.

Layer 2 and Layer 3 designs influence failover. Interface monitoring, link aggregation, dynamic routing, virtual MAC behavior, upstream ARP or neighbor caches, and switch convergence can all affect how quickly traffic reaches the new primary. The firewall may complete failover before the rest of the network is ready to use it.

Architecture diagrams should therefore show dependencies outside the cluster. The useful question is not whether the HA status page is green; it is whether an application session still has a complete path after a specified component fails.

Session pickup should be used according to application need

FortiGate clustering can synchronize TCP session state so a new primary can resume many existing sessions after failover. That reduces interruption, but synchronization consumes resources and does not guarantee that every protocol will continue seamlessly under every condition.

Connectionless traffic, NAT state, expectation sessions, IPsec tunnels, and application behavior may require additional consideration. Some applications recover cleanly by reconnecting, while long-lived transactional sessions may experience user-visible failure.

Teams should identify which traffic genuinely needs state preservation. Enabling every synchronization option without understanding the value can add overhead while creating unrealistic expectations about zero-impact failover.

Routing convergence is part of firewall availability

Dynamic routing can make a firewall highly available across Layer 3 boundaries, but it adds timers and protocol state to recovery. When a primary changes or an entire site disappears, neighbors must detect the change, withdraw or replace routes, and select a viable path.

Fast detection technologies and tuned routing timers can reduce outage duration, yet aggressive settings can cause instability on noisy links. The correct design balances convergence speed with the quality of the underlying network.

Routing should also be tested during planned maintenance. A design that only works when a device crashes may behave differently when engineers gracefully remove it from service or change route advertisements.

Capacity planning must assume the failure state

Redundant systems often run comfortably during normal operation because traffic is distributed or the primary handles a predictable load. After a failure, surviving components may carry more traffic, more encrypted tunnels, more logging, and more inspection work.

Each surviving firewall, circuit, and upstream device should have enough headroom for the expected failure case. That may not require full duplication of peak capacity for every environment, but the trade-off should be explicit. Otherwise the failover succeeds technically and then degrades under load.

PrepAway’s comparison of high availability and fault tolerance uses a cloud context, but the conceptual distinction applies here too: adding redundancy does not automatically eliminate every interruption or shared dependency.

Capacity also includes control-plane resources. A failover can trigger route recalculation, VPN renegotiation, log bursts, and management activity at the same time that the surviving firewall is processing more user traffic. CPU, memory, session tables, and cryptographic throughput should be observed during failure testing rather than inferred only from steady-state utilization.

Disaster recovery requires independent dependencies where they matter

A secondary site is not a true recovery option if it depends on the same carrier, identity service, management plane, DNS platform, certificate authority, or upstream network that failed with the primary site. Some shared services can be intentionally global, but their resilience needs to be included in the design.

Public services may require alternate internet connectivity, route advertisement, load balancing, or DNS changes. Site-to-site VPNs may need peer configurations that allow a remote office to reach either site. Management access must work even when the primary control path is unavailable.

PrepAway’s disaster recovery planning coverage provides broader continuity context for defining recovery priorities, dependencies, and testing rather than focusing on one network device.

Configuration recovery must be protected from bad changes

Hardware failure is visible. Configuration failure can be more subtle and can replicate across redundant members. If an incorrect policy, route, object, or certificate is synchronized to the cluster, HA preserves the mistake rather than protecting against it.

Backups, version control, staged deployment, approval workflows, and known-good restore procedures are therefore part of resilience. Teams should be able to recover not only from a failed appliance but also from a harmful change made by an authorized administrator.

Centralized management can help by preserving revision history and controlling deployment, but it also becomes a dependency. Recovery procedures should explain how critical changes are made if the normal management platform is unavailable.

Recovery data should be stored outside the failure domain it is intended to recover. A configuration backup on a management server located in the same data center as the protected firewall may be unavailable during the event that requires it. Critical licenses, certificates, bootstrap configuration, contact information, and recovery instructions should be accessible through an alternate operational path.

Testing should include degraded and messy failures

Clean power-off tests are useful but incomplete. Real incidents include packet loss, half-open links, failed routing adjacencies, saturated circuits, broken synchronization, partial site isolation, expired certificates, and upstream provider problems. These conditions can create split-brain or asymmetric behavior that is harder than a complete outage.

A mature test plan exercises the failures the architecture claims to survive and records actual recovery time. It should verify application traffic, management access, monitoring, logs, VPNs, and the ability to fail back after the original problem is resolved.

Fortinet-specific operational scenarios are explored in PrepAway’s advanced troubleshooting and high-availability article, which is useful background for understanding why resilience needs hands-on validation.

Resilience is an operational contract with the business

Network teams should be able to describe what the firewall service will do during defined failures. Statements such as ‘we have HA’ are too vague. Better commitments specify whether sessions may reset, how long routing convergence is expected to take, which site becomes active during disaster recovery, and which dependencies could extend recovery.

Those commitments should match application requirements. A customer-facing payment service, internal file server, and development environment may need different recovery objectives. Designing every system for the strictest requirement can be expensive, while applying one weak standard everywhere can leave critical services exposed.

The broader Fortinet certifications cover the technologies used to build these designs, but resilient architecture ultimately depends on matching failure behavior to business impact and then proving the design through repeatable testing.

Recovery objectives should also be revisited after major architecture changes. Moving applications to cloud platforms, changing internet providers, consolidating data centers, or adopting new remote-access patterns can create dependencies that did not exist when the HA design was approved. A resilience review after these changes helps ensure that the documented failover path still matches the path real traffic will take during an outage.

Related Posts

• PKI in Practice: Certificates, Trust Chains, and Failure Modes

• Vulnerability Management Beyond the Scanner

• Managed Identities: Stop Treating Credentials as Application Configuration

• How Routers Really Decide Where Packets Go

• Identity Is the New Security Perimeter

• Troubleshooting Layer 2 Before Blaming Layer 3

• Zero Trust Is a Design Principle, Not a Product

• Foundation Model Choice Is a Product Decision as Much as a Technical One

• OSPF at Enterprise Scale

• NETCONF, RESTCONF, or APIs?