Palo Alto Networks NGFW-Engineer: High Availability on Palo Alto Firewalls
High availability on Palo Alto firewalls is not just a checkbox that creates a redundant appliance. An HA design has to synchronize the configuration and session state that should survive a failure, detect when the active path is no longer healthy, move traffic to the peer, and integrate with the surrounding switches, routers, VPNs, and applications. A pair can report healthy HA status and still fail to provide useful service if the upstream or downstream network does not follow the transition.
The current NGFW Engineer scope treats management and operation as part of firewall engineering for that reason. HA is a system property built from peer communication, synchronized state, monitored paths, election behavior, and network convergence.
Choose active/passive unless active/active solves a real requirement
In active/passive mode, one firewall normally processes traffic while the peer remains ready to take over. The model is straightforward because there is one active forwarding path at a time. Active/active allows both peers to pass traffic and can support designs that need that behavior, but it introduces additional complexity in session ownership, routing, and surrounding network design.
The right choice depends on the platform and use case; some newer hardware families support only active/passive. Do not choose active/active merely because it sounds like better utilization. The operational cost of a more complex topology should be justified by a requirement that active/passive cannot satisfy.
HA1 carries control-plane coordination
The HA1 control link is used for peer control communication, including state information and configuration-related coordination. Palo Alto Networks supports a primary and optional backup control link, and heartbeat backup can use the management path to help prevent split-brain conditions when the primary control link is lost.
A split-brain scenario occurs when both peers believe the other has failed and each attempts to become active. Redundant control communication reduces that risk. The HA control path should therefore be treated as production infrastructure, with appropriate interfaces, addressing, monitoring, and protection rather than an incidental cable between appliances.
HA2 synchronizes dataplane state
The HA2 data link carries dataplane synchronization such as the session table and other runtime state. In active/passive operation, session synchronization allows the passive firewall to receive state created by the active peer so it can continue many sessions after failover instead of rebuilding every connection from scratch.
Palo Alto Networks documentation identifies the session table as HA2-synchronized data. The data link can also have a backup path, and HA2 keep-alive monitoring can detect failures beyond a simple physical link-down event. Engineers should verify actual HA2 counters and synchronization status during testing rather than assuming the link is healthy because the interface is up.
Define failover conditions around service health
A firewall should fail over when the active peer can no longer provide the intended service, not only when the appliance loses power. Link monitoring can watch important physical interfaces. Path monitoring can test reachability to selected destinations through virtual routers or other forwarding contexts. Those mechanisms allow HA to react when the firewall is running but disconnected from a critical part of the network.
Monitoring needs careful scope. Failing over on one noncritical interface can cause unnecessary instability, while monitoring too little can leave a broken active firewall in service. Group interfaces and path targets according to the business service they represent and test the difference between “any” and “all” failure conditions.
Device priority and preemption affect which peer becomes active
HA peers use election settings to determine the preferred active device. Device priority can express preference, and preemption can allow a higher-priority peer to reclaim the active role after it returns to service. Preemption is not always desirable because a second transition can interrupt traffic again after the environment has already stabilized on the surviving peer.
Choose the behavior around operational policy. Some organizations prefer a deterministic primary device; others prefer to leave the healthy peer active until a planned maintenance window. The key is to make the decision explicit and include it in upgrade and incident procedures.
Session synchronization has limits
Not every session or runtime object is synchronized in every HA mode. Palo Alto Networks documents several exceptions, and decrypted SSL sessions are a particularly important example because proxy-decrypted sessions are not synchronized between peers. A failover can therefore cause application reconnection even when normal session synchronization is working.
The decryption policy article highlights this limitation. HA testing should include encrypted applications that matter to the business so teams understand which sessions survive and which users will see reconnect behavior.
The surrounding network must converge with the firewall
HA state alone does not move packets through external switches and routers. Interface state, virtual MAC behavior, dynamic routing, LACP, upstream ARP or neighbor information, and cloud networking constructs can all influence how quickly traffic reaches the newly active peer. The firewall pair and the adjacent network should be tested as one failover domain.
This is why system-level availability is a useful broader principle. Redundant appliances are only one part of availability. The path to the application, authentication services, DNS, logging, and management must also tolerate the event.
Plan upgrades as controlled failovers
HA allows software upgrades and maintenance to be staged, but the process still requires discipline. Verify synchronization, move traffic intentionally, confirm the surviving peer is healthy, upgrade the passive or non-forwarding device as appropriate, return it to functional state, and repeat for the other peer. Product-version compatibility rules should be followed during the transition.
After each step, confirm real traffic, not only HA widgets. Session counts, application tests, routing, VPNs, and logging should be checked before proceeding. A maintenance window is not the time to discover that path monitoring or upstream convergence was never tested.
Test failure modes that operators will actually face
A useful HA test plan includes device failure, loss of monitored interfaces, loss of a routed path, HA1 disruption, HA2 disruption, upstream circuit failure, and maintenance-driven failover. Each test should record detection time, state transition, network convergence, session impact, and application recovery.
HA testing should distinguish control-plane failover from service recovery. A peer can become active quickly while upstream routing, ARP or neighbor state, VPNs, dynamic protocols, or application dependencies take longer to converge. Measure the user-visible interruption for representative flows rather than only the time between HA state transitions. The traffic troubleshooting workflow can verify where packets stop during that interval and whether the new active firewall has the routes, policy, NAT, and identity context required to forward them.
Maintenance design is equally important. Teams should document which peer is upgraded first, how configuration synchronization is verified, when preemption is disabled or restored, and what condition aborts the maintenance window. If the environment uses Panorama, pushes should be sequenced so that both peers do not receive an unvalidated network change immediately before a controlled failover. The broader network security platform remains available only when firewall HA and management change control support each other.
Capacity planning must assume the surviving peer carries the full protected workload. CPU, session tables, content inspection, decryption, tunnel counts, routing scale, and interface utilization should all remain within safe margins after failover. This is particularly important for encrypted traffic because some decryption state has synchronization limitations. Candidates preparing for the NGFW Engineer exam should treat redundancy as a complete service-capacity problem, not merely a pair-status configuration.
Split-brain prevention deserves explicit testing as well. HA links, heartbeat backup, management reachability, and election settings exist to help peers distinguish a failed partner from a failed communication path. A design that has redundant firewalls but a single fragile HA transport can create ambiguous states during a network failure. Test loss of HA1, HA2, monitored links, and monitored paths separately so operators know which failures cause takeover, which degrade synchronization, and which require manual intervention. The goal is deterministic behavior, not merely more cables between two appliances.
Runbooks should capture the expected steady state after recovery as well as the failover action. Once service is restored, verify synchronization, routing adjacencies, monitored paths, interface state, and whether the preferred peer should resume the active role. Recovery that leaves the pair in an undocumented state simply postpones the next failure.
The cross-platform FortiGate HA discussion is relevant because the engineering lesson is vendor-independent: failover is only proven when the service continues under realistic failure conditions. Repeating tests after network or policy changes prevents old assumptions from becoming part of the runbook.
High availability on Palo Alto firewalls works when peer synchronization, failover detection, election behavior, and the external network are designed together. HA1 keeps the peers coordinated, HA2 carries critical dataplane state, and link or path monitoring tells the pair when the active path is no longer useful.
The final measure is application continuity. If a controlled failure produces the expected active peer, predictable session impact, correct routing, and a recoverable operating state, the HA design is doing its job. If operators only know that both firewalls display “healthy,” the most important part of the design has not yet been tested.