Why Service Providers Keep Choosing IS-IS
IS-IS has survived repeated predictions that it would be displaced because it fits the shape of large provider networks unusually well. It is a link-state IGP with a compact hierarchy, flexible TLV-based extensions, and a history of carrying new control-plane information without forcing a redesign of the basic protocol. The current 350-501 SPCOR exam still expects engineers to implement and troubleshoot IS-IS for IPv4 and IPv6 as part of the CCNP Service Provider certification.
The reason operators keep choosing it is not nostalgia. A provider needs an IGP that converges predictably, carries infrastructure reachability, scales across a large topology, and can absorb extensions such as segment routing. IS-IS does those jobs while keeping the core routing domain conceptually separate from IP transport because the protocol itself runs directly over Layer 2 rather than using IP packets for neighbor formation.
That architecture can feel unfamiliar to engineers who first learned OSPF, but the operational questions are familiar: who is my neighbor, what link-state information have I learned, what SPF result should follow, and which routes should enter the RIB?
IS-IS separates routing hierarchy into Level 1 and Level 2
IS-IS uses Level 1 for intra-area routing and Level 2 for inter-area routing. A router can participate in one level or both, and Level 1-2 routers connect the local area to the Level 2 backbone. This creates a hierarchy that can limit topology scope without requiring every router to carry every detail from every area.
The hierarchy is a design tool, not a reason to create many areas by default. Too many boundaries add route leaking and operational complexity; too few can increase link-state database and SPF impact. The same trade-off appears in broader network design: hierarchy is valuable when it contains failure and state, not when it is added simply because the protocol supports it.
NET addressing looks unusual but follows a clear structure
An IS-IS Network Entity Title includes an area address, a system ID, and an NSEL. The system ID uniquely identifies the router inside the IS-IS domain, while the area portion determines Level 1 membership. The address is not used like an IPv4 interface address, which is why engineers can initially overthink it.
Operationally, the important rules are uniqueness and consistency. Duplicate system IDs create serious ambiguity, and inconsistent area planning produces unexpected adjacency or route-leaking behavior. Once a deterministic scheme is chosen, NET addressing becomes background infrastructure rather than a daily troubleshooting mystery.
Link-state PDUs and TLVs make IS-IS extensible
IS-IS advertises topology information in link-state PDUs. Information is carried in type-length-value structures, which means new capabilities can be added through new TLVs without changing the core packet architecture. That extensibility is one reason IS-IS adapted well to IPv6, traffic engineering, and segment routing.
For an engineer, TLVs matter because they connect protocol state to features. A segment-routing SID, a metric, an IPv6 reachability item, or a capability is not abstract magic; it is information advertised into the link-state database. The 350-501 SPCOR core technologies are useful context because they show how the IGP becomes the substrate for MPLS, SR, assurance, and services.
Wide metrics are essential for modern provider engineering
Modern IS-IS designs use wide metrics because they support larger values and the extended information required by traffic engineering and segment routing. Metric planning should still be intentional. The IGP metric expresses a provider’s default topology preference, and excessive manual tuning can make failure behavior difficult to predict.
A good metric design reflects topology and capacity at a useful level, then leaves specialized service steering to policies designed for that purpose. If every traffic-engineering requirement is solved by changing IGP cost, the underlay becomes a collection of hidden exceptions and the shortest-path model loses its explanatory value.
IS-IS is a natural carrier for segment-routing information
Segment routing strengthens the case for a clean IS-IS domain because the IGP advertises prefix and adjacency SIDs in addition to reachability. In SRv6 designs, IS-IS can advertise locators and SRv6 capabilities. This allows the same link-state topology to support shortest-path forwarding, explicit segment lists, fast reroute, and flexible algorithms.
The combination rewards engineers who understand both conventional routing and modern path programming. A router that loses an IS-IS adjacency may simultaneously lose IP reachability and the segment information that higher-level SR policies depend on. Advanced routing knowledge from enterprise routing troubleshooting transfers well here even though the provider scale and service consequences are different.
Fast convergence depends on timers, detection, SPF, and repair together
Reducing a hello timer alone does not guarantee a fast or stable provider network. Failure detection, LSP generation, flooding, SPF scheduling, RIB/FIB programming, and any local repair mechanism all contribute to convergence. Aggressive timers can create churn if links flap or CPU resources are constrained.
A better design defines a convergence objective, uses mechanisms such as BFD where appropriate, and validates the whole failure sequence under realistic load. This is why network troubleshooting needs to be more than checking whether a neighbor is up: a routing system can have all neighbors restored while traffic still follows a suboptimal or temporarily inconsistent path.
Troubleshooting IS-IS starts with adjacency, database, SPF, and RIB
A disciplined sequence keeps investigations small. First prove that the expected neighbors are present at the correct levels. Then confirm that the required LSP is in the database and current. Next check the computed route and metric, followed by RIB and forwarding installation. If segment routing is involved, verify the relevant SID advertisements and programmed label or SRv6 state.
This order prevents random configuration changes. If the LSP is absent, changing a route policy will not help. If the database is correct but the route is not installed, the problem is downstream of flooding. If the route is installed but traffic fails, the investigation should move to forwarding, MPLS/SR state, or the service layer.
IS-IS lasts because it keeps the provider core understandable
Provider networks are complex enough without an IGP that needs constant reinvention. IS-IS gives operators a stable link-state model, a scalable hierarchy, and extensibility through TLVs. It also coexists well with BGP at the service edge and segment routing in the transport core, which is why it remains common even as architectures change. The broader network-architecture perspective reinforces the principle: durable protocols survive when they provide clear abstractions and predictable failure behavior.
Choosing IS-IS does not eliminate design work. Area structure, metrics, summarization or leaking, convergence tuning, and extension support still need engineering. But the protocol gives those decisions a coherent place to live. Service providers keep choosing it because its fundamentals remain stable while its information model continues to accommodate the next transport and automation requirement.
Provider engineers should also understand how IS-IS flooding behaves under stress. Link-state information must reach the routers that need it quickly, but excessive churn can consume control-plane resources. LSP generation, pacing, SPF scheduling, and flooding controls are therefore part of stability engineering. Tuning should be based on measured convergence goals and device capabilities rather than copied timer values. A network that converges extremely fast in a small lab may behave differently when thousands of prefixes and many simultaneous adjacency events are present.
Level boundaries deserve periodic review because topology changes can make an old hierarchy inefficient. A new metro ring, backbone merge, or acquisition may leave traffic repeatedly crossing Level 1-2 boundaries or create accidental dependence on default routing where explicit reachability would be clearer. The right hierarchy is not fixed forever. It should reflect the current failure domains and operational ownership while preserving the simplicity that made IS-IS attractive in the first place.
When IPv4 and IPv6 share IS-IS, single-topology and multi-topology choices affect whether both address families are assumed to have the same connectivity. If all links support both families, a shared topology can simplify operations. If some links or services differ, separate topology information may be necessary. The engineer should know which model is active before interpreting a route absence, because the protocol can be healthy while one address family legitimately follows a different topology.
Finally, IS-IS should be monitored as a system rather than as a neighbor count. Database size, LSP age, sequence changes, SPF frequency, overload state, adjacency churn, and route installation all reveal different kinds of stress. A router can maintain every adjacency while repeatedly regenerating LSPs because of an unstable interface attribute. Continuous visibility into those patterns allows teams to fix a degrading control plane before it becomes a customer-visible routing event.
Maintenance procedures are another reason IS-IS remains attractive in large cores. Overload-bit behavior can keep a router participating in the control plane while discouraging it from carrying transit traffic, which is useful during staged maintenance or resource stress. Used carefully, this gives operators a way to drain forwarding without tearing down every adjacency at once. The same feature can also mask deeper problems if it is left enabled accidentally, so operational checks should treat overload state as intentional configuration that needs an owner and a reason.
A final design review should also confirm that operational teams can distinguish ordinary Level 1 or Level 2 behavior from redistribution or leaking that was added as an exception. Clear route ownership makes future troubleshooting faster and prevents a temporary workaround from becoming an invisible dependency in the provider core.