Practice Exams:

Diagnosing Enterprise Routing Failures

 

Routing failures are rarely solved by memorizing the largest number of show commands. The fastest engineers reduce the problem in layers. They define the symptom, determine its scope, establish whether the failure is control plane or forwarding plane, and then follow the route from source to destination until the expected state disappears. Each observation narrows the search.

That method connects directly to the current 350-401 ENCOR expectation that engineers diagnose network problems with tools such as ping, traceroute, debugs, SNMP, and syslog. The 300-410 ENARSI path extends the same discipline into more advanced routing and services troubleshooting.

The goal is not to make the network “green.” It is to explain why the forwarding decision differs from the intended path and prove the correction at every relevant layer.

Start with a precise symptom

“Routing is broken” is not actionable. Identify source, destination, protocol, direction, first observed time, affected sites, and whether the failure is complete or intermittent. A single application failing while other destinations work suggests a different problem from an entire branch losing every remote prefix.

Confirm the symptom from more than one vantage point if possible. Test from the affected host or a nearby gateway, then compare with another source. A failure limited to one VRF, VLAN, address family, or source prefix can quickly narrow the routing context.

Do not begin by changing timers or clearing protocols. Preserve evidence first, especially when the problem is intermittent.

Users often report a “network” problem when DNS, authentication, a firewall rule, or the application itself is failing. Test the destination by address when appropriate and compare transport behavior. If the IP path works but the application name fails, routing may be healthy.

Similarly, successful ping does not prove the application path is healthy. ICMP can take a different policy path, and firewalls may treat it differently. Use tests that resemble the affected traffic while staying within operational policy.

This early separation prevents the routing team from spending an hour debugging OSPF for a DNS outage.

Verify the source has the expected forwarding decision

On the first Layer 3 hop, check the route that should carry the destination. Is the prefix present? Which protocol installed it? What administrative distance and metric won? What next hop and outgoing interface are selected?

If the prefix is absent, the problem is control-plane learning, filtering, summarization, redistribution, or route generation. If the prefix is present but points somewhere unexpected, investigate best-path logic or policy. If the route is correct, move toward next-hop reachability and the forwarding table.

Advanced context from ENARSI routing is useful because the route installed in the RIB is the result of several prior decisions, not a raw copy of every advertisement the router has heard.

Adjacency health is necessary, not sufficient

An OSPF or BGP neighbor can be established while the required routes are missing. Verify the adjacency state, then verify what is being advertised, received, accepted, and installed. A healthy session only proves the peers can maintain the protocol relationship.

For OSPF, investigate area scope, LSAs, network type, filtering, summarization, and whether the expected prefix is actually originated. For BGP, examine address family, policy, best-path selection, next-hop reachability, and inbound or outbound filtering.

Treat “neighbor is up” as one checkpoint in the chain, not as proof that routing is correct.

Recursive next hops can hide the real failure

A route may point to a next-hop address that itself requires resolution through another route. If recursion resolves through the wrong interface, disappears after a summary change, or depends on a route that is filtered, the top-level route can look present while forwarding fails.

Follow the recursion until the router reaches a directly usable adjacency. Verify ARP or ND where appropriate and confirm the outgoing interface is operational. On platforms with CEF or equivalent forwarding structures, compare the RIB decision with the FIB entry.

This distinction matters because control-plane state can lag or diverge from the actual forwarding programming during software defects, hardware issues, or convergence.

Traceroute can show where replies stop or paths diverge, but missing hops do not always mean forwarding stops there. Devices may rate-limit TTL-expired responses, filter probes, or return traffic asymmetrically. Interpret the pattern alongside routing tables and interface state.

A sudden path change can still be highly informative. Compare a working source with a failing source, or compare before and after an incident if historical path data exists. The first meaningful divergence suggests where to inspect policy or reachability next.

Active path evidence is strongest when paired with device state rather than treated as a complete diagnosis.

Policy can override the route you expected

Policy-based routing, route maps, VRF leaking, SD-WAN policy, security service insertion, and firewalls can alter traffic independently of ordinary destination lookup. A route table that looks perfect may not represent the actual packet decision.

Identify which policies apply to the ingress interface or flow. Verify match counters and default behavior. In complex enterprise designs, document the order of operations so the team knows whether classification occurs before or after a particular routing decision.

The 300-420 ENSLD design perspective helps because troubleshooting becomes easier when the architecture makes policy boundaries explicit rather than scattering exceptions across devices.

Asymmetry changes what “working” means

Forward traffic can reach the destination while the return path fails or crosses a stateful device that never saw the original flow. This produces symptoms such as successful one-way UDP, partial TCP handshakes, or application resets that look inconsistent.

Check both directions. On multihomed sites, BGP policy, ECMP, first-hop redundancy, or default-route changes can create a return path different from the forward path. Stateful firewalls and NAT make that asymmetry especially important.

Routing troubleshooting should therefore trace the conversation, not merely the outbound packet.

Logs and telemetry give you the time dimension

A route table shows the network now. Many incidents depend on what changed 20 minutes ago. Syslog, routing event logs, interface transitions, telemetry, and configuration history can reveal the sequence that produced the current state.

Look for adjacency flaps, interface errors, policy deployment, CPU pressure, or route-count changes around the first observed failure. Correlation is not proof, but it gives a hypothesis to test.

This is where modern assurance complements classic CLI troubleshooting. Historical evidence can preserve a transient condition that vanished before an engineer reached the console.

Fix the cause and verify the service

Once the likely cause is found, make the smallest correction that restores intended behavior. Do not use broad route clearing or configuration rollback as the first move unless risk demands it; those actions can erase evidence or trigger a larger convergence event.

After the fix, verify the control plane, RIB, FIB, next-hop reachability, and end-to-end service. Confirm related sites did not lose reachability and that the route remains stable long enough to consider the incident resolved.

The broader CCNP Enterprise skill is repeatability: an engineer should be able to explain exactly which expected state was missing, why it was missing, and what evidence proves the network now matches the intended design.

A troubleshooting method becomes an operational asset

After the incident, convert the reasoning into monitoring or automation. If a missing BGP community caused the outage, add validation for that community. If a recursive next hop failed silently, monitor the dependency. If an undocumented PBR rule surprised the team, document and test policy scope.

Organizations improve when each failure makes the next similar failure easier to isolate. Runbooks should capture decision points and evidence, not just a list of commands. “If prefix absent here, inspect these three causes” is more valuable than twenty pages of show output.

Routing failures stop feeling mysterious when engineers follow state in order: symptom, scope, route learning, path selection, recursion, forwarding, policy, return path, and service verification. The network usually tells you where the failure is; the discipline is to ask one layer at a time.

Some of the hardest routing failures occur where information crosses a boundary: one protocol into another, one VRF into another, or one administrative domain into another. Redistribution can change metrics, tags, route types, and loop-prevention behavior. VRF leaking can create reachability that exists in one table and is absent in another.

At those points, verify both sides of the boundary explicitly. Confirm the source route exists before export, that policy permits it, that attributes are set as intended, and that the receiving process actually installs it. Do the same for the return direction. A summary or default route can hide the absence of a more specific route until traffic reaches the boundary and fails.

Boundary troubleshooting also benefits from route tagging and clear policy names. When the route carries evidence of where it was redistributed or which policy handled it, engineers can reconstruct the path faster than when every redistribution point uses anonymous route maps and undocumented defaults.

Packet captures can settle disputes when control-plane views disagree with observed behavior. A capture on the ingress and egress sides of a suspected boundary can show whether the packet arrived, whether the next hop changed, whether TTL or DSCP values were altered, and whether return traffic exists. Use captures selectively; they are evidence, not a substitute for understanding the forwarding decision.

During major incidents, keep a timeline of hypotheses and observations. Record when a route disappeared, when a policy was changed, when a neighbor recovered, and when user traffic returned. That timeline prevents teams from repeating disproven theories and gives the post-incident review a factual sequence instead of memories reconstructed after the pressure is over.

Related Posts

• How Attack Paths Form Across Enterprise Systems

• Start With Risk When Choosing Security Controls

• Azure RBAC: Separate Scope From Role

• Azure Backup and Site Recovery Protect Against Different Failures

• Subnetting Gets Easier When You Stop Memorizing Tables

• DHCP and DNS: Two Services That Make Everything Else Look Broken

• REST APIs for Network Engineers Who Grew Up on the CLI

• Designing an Enterprise Core for Failure

• Serverless Still Needs Capacity Planning

• From Monolith to AWS: Choose the Migration Pattern That Fits