FortiGate High Availability: What Fails Over and What Does Not
High availability is often described as if two firewalls simply become one reliable firewall. FortiGate clustering is more precise than that. FGCP can synchronize configuration and state across cluster members and can move the active role when a monitored failure occurs, but not every condition is detected automatically and not every type of traffic state is guaranteed to survive in the same way.
High availability is explicitly part of the administrative work covered by the current FortiOS 7.6 Administrator exam. The important skill is not reciting active-passive versus active-active. It is knowing what the cluster monitors, what it synchronizes, what external dependencies remain, and what users experience when failover actually happens.
Within the broader Fortinet certification portfolio, HA becomes part of larger resilient-network and operations designs. At the administrator level, a useful principle is that redundancy only protects the components and failure modes the architecture has intentionally covered.
A pair of FortiGates cannot compensate for a single upstream switch, a shared power circuit, an unsynchronized routing dependency, or an application session that cannot recover. HA should therefore be tested as a system, not celebrated because the dashboard says the cluster is healthy.
FGCP creates a cluster with one shared operating identity
FortiGate Clustering Protocol allows compatible FortiGate units to form a cluster. Members must meet platform and firmware requirements, exchange heartbeat information, synchronize configuration, and elect roles. To the surrounding network, the cluster is designed to behave like a highly available firewall service rather than two independent devices.
That shared identity simplifies operations, but administrators still need to know which settings are synchronized and which remain member-specific. Management addressing, priorities, some interface details, and diagnostic context can differ. Treating every field as globally shared makes troubleshooting harder when a standby member behaves differently after promotion.
Heartbeat health is the foundation of cluster decisions
Cluster members need reliable heartbeat communication to know whether peers are present and healthy. Heartbeat design should therefore be physically and logically resilient enough that a single cable or switch failure does not cause the cluster to misinterpret peer status.
Split-brain conditions are especially dangerous because two devices can believe they should be active. The HA design should reduce that risk through appropriate heartbeat interfaces, cabling, and topology. Redundancy that introduces ambiguous ownership is worse than a clearly failed standalone device.
Monitored interfaces determine which path failures trigger failover
An active firewall can remain powered on while one critical link fails. Interface monitoring lets the cluster treat selected connectivity failures as reasons to move the active role. The design question is which interfaces are truly service-critical and whether monitoring them reflects end-to-end availability.
Monitoring too little leaves the cluster active while an important path is broken. Monitoring too much can create unnecessary failovers for interfaces that do not justify a role change. The policy should follow business dependency rather than simply marking every port important.
Configuration synchronization is necessary but not the same as service continuity
FGCP synchronizes most operating configuration so a secondary member can take over with the same policies, routes, objects, VPN settings, and security controls. That solves one class of failover problem: the standby does not need to be reconfigured during the incident.
Service continuity also depends on neighboring devices accepting the new forwarding behavior, dynamic protocols converging where applicable, ARP or neighbor state updating, and traffic reaching the active member. A synchronized configuration can still sit behind an upstream network that does not recover cleanly.
Session synchronization determines what users feel
Whether an established connection survives failover depends on what session state has been synchronized and what the application can tolerate. TCP, UDP, NAT, expectation sessions, and IPsec state do not all have identical behavior under every HA mode and configuration.
This is where a design should distinguish “the firewall failed over” from “the application session survived.” A user may see a reconnect even though the cluster behaved exactly as designed. If uninterrupted sessions are a requirement, the HA test must measure session continuity rather than only ping reachability.
Active-active does not eliminate the need for an active control role
FortiGate active-active clustering can distribute some processing, but it should not be interpreted as two independent firewalls both making unrelated control decisions. The cluster still has coordinated roles and shared configuration, and traffic distribution has platform-specific behavior.
Administrators should select HA mode based on capacity, failure behavior, and operational simplicity. Active-active is not automatically “more redundant” than active-passive. A simpler active-passive design may be easier to validate and support when throughput requirements do not justify additional distribution complexity.
External dependencies are part of the failover domain
A firewall pair can share a single switch, ISP handoff, authentication source, DNS dependency, log collector, or power path. Those shared components can defeat the purpose of the cluster if they fail independently. A complete HA review maps dependencies on both sides of the FortiGate rather than stopping at the two appliances.
Historical material on Fortinet high availability and SD-WAN troubleshooting illustrates why resilient designs need end-to-end thinking. The exact exam generation is older, but the architectural lesson remains useful: redundancy must include path and state, not just duplicate hardware.
Planned failover tests should be more demanding than pulling power
A useful test program exercises different failure modes: device power loss, monitored-interface failure, heartbeat disruption, upstream path loss, session continuity, VPN recovery, routing convergence, and restoration to normal service. Each test should have an expected outcome and measurable recovery behavior.
That evidence also exposes hidden operational assumptions. If a failover succeeds only when an engineer manually clears an upstream table, the design is not truly automatic. If a business-critical application requires a fresh login, that may be acceptable—but it should be known before a real outage.
High availability is successful when failure behavior is predictable
The objective of HA is not to claim zero downtime in the abstract. It is to make defined failures produce defined recovery behavior within an acceptable interruption window. That requires synchronized configuration, appropriate state handling, healthy heartbeat links, monitored dependencies, and a surrounding network that recognizes the active member.
Document what does not fail over as carefully as what does. Member-specific management details, unsynchronized state, external services, and application behavior can all influence recovery. Those limits are part of the architecture and should be visible to operators.
A mature FortiGate HA design can answer three questions before an outage: what failure will the cluster detect, what state will move to the surviving unit, and what will users notice? If those answers are tested rather than assumed, high availability becomes an operational capability instead of a checkbox on an architecture diagram.
Firmware lifecycle deserves special treatment in a cluster because HA does not remove upgrade risk. Administrators need a supported upgrade path, configuration backup, compatibility review, and a plan for how members will transition while maintaining service. The fact that a secondary member exists is helpful, but it does not make an unsupported or poorly tested firmware change safe.
Routing protocols can also alter the user-visible recovery time. If the FortiGate participates in BGP or OSPF, neighbors may need to reconverge or re-establish adjacencies after a role change depending on topology and feature behavior. The cluster can become active quickly while the surrounding network still needs time to restore the best path. Measure end-to-end application recovery, not only HA election time.
VPNs deserve their own test cases. Site-to-site tunnels, remote-access sessions, and overlay networks carry state that may interact with session synchronization, peer detection, and routing. A failover test should verify whether tunnels remain established, renegotiate automatically, or require the application to reconnect. Different VPN designs can have different acceptable interruption windows.
Management access should remain possible during degraded states. Reserved management interfaces, out-of-band paths, and member-specific access can help an operator diagnose the unit that lost the active role without disrupting production. If the only management path depends on the same network that failed, the cluster may preserve user traffic while leaving administrators blind.
Capacity planning must assume the surviving member can carry the required workload. Two appliances running comfortably at 55 percent each in active-active mode may not provide useful redundancy if one remaining unit cannot sustain the combined inspection load after failure. HA capacity is therefore a worst-case design problem. Measure CPU, memory, session scale, and security-processing headroom under the load the survivor is expected to inherit.
Stateful security services should be included in resilience testing too. If a cluster uses external authentication, FortiGuard services, certificate validation, or central logging, the surviving unit still depends on those services after failover. A firewall pair can remain healthy while users lose authentication or security visibility because a shared external dependency failed at the same time.
Failback deserves as much attention as failover. Returning service to the preferred member can trigger another state transition, route update, or session disruption. Some organizations choose controlled failback during a maintenance window rather than automatically moving back as soon as the original unit recovers. The correct choice depends on whether stability or rapid restoration of the preferred topology is more important.
Operational documentation should identify which unit owns the active role, how priorities are configured, which interfaces are monitored, where management access resides, and which external systems need to be checked during a failover. That runbook turns a cluster event into a known procedure instead of forcing responders to rediscover the HA design while production traffic is already at risk.
After every planned test, record measured interruption rather than a simple pass or fail. Recovery time, lost sessions, routing convergence, and management visibility provide a better baseline for future changes.