Practice Exams:

VXLAN EVPN Is Easier When You Separate the Overlay From the Underlay

 

VXLAN EVPN can feel complicated because several control and data-plane ideas are discussed at once: leaf-spine routing, VTEPs, VXLAN tunnels, VNIs, BGP EVPN routes, anycast gateways, endpoint learning, and external connectivity. The current 350-601 DCCOR exam and CCNP Data Center certification expect engineers to understand data-center network technologies as systems, and the cleanest mental model is to separate the IP underlay from the VXLAN EVPN overlay.

The underlay answers a transport question: can one VTEP reach another reliably across the routed fabric? The overlay answers a tenant-connectivity question: which endpoints, networks, and routes exist, and how should traffic be encapsulated and forwarded between VTEPs? Problems become easier to reason about when an engineer first decides which of those layers is broken.

That separation does not mean the layers are independent in operation. The overlay depends on underlay reachability, MTU, convergence, and stable loopbacks. It means each layer has a distinct job that can be tested before moving to the next.

The underlay is a routed IP transport

The fundamentals behind IP addressing and subnetting still matter in a modern fabric. Leaf and spine nodes establish routed adjacencies, advertise loopback reachability, and create multiple equal-cost paths across the topology. The underlay does not need to know tenant MAC addresses; it needs to deliver encapsulated packets between tunnel endpoints.

Cisco documentation describes OSPF, IS-IS, and eBGP as underlay options in Nexus VXLAN fabrics. The engineering criteria include fast convergence, simplicity, addressing design, scale, and operational familiarity. Whatever protocol is chosen, VTEP reachability must be stable before overlay troubleshooting begins.

The overlay carries tenant reachability and segmentation

VXLAN creates logical segments identified by VNIs and encapsulates tenant frames or routed traffic for transport across the IP fabric. EVPN provides the control plane that distributes reachability instead of relying entirely on flood-and-learn behavior. Leaf switches acting as VTEPs can therefore learn where remote endpoints or subnets reside.

This separation lets workloads keep logical network membership without requiring the physical fabric to extend large Layer 2 domains everywhere. The underlay stays routed while the overlay supplies virtualization and tenant boundaries.

Leaf-spine topology makes path behavior predictable

Network architecture becomes easier to operate when the topology has a regular structure. In a leaf-spine fabric, leaf switches connect endpoints and every leaf has paths through the spine layer to other leaves. Equal-cost multipath can use that symmetry to spread flows and avoid a single hierarchical aggregation path.

Spines provide transit between leaves; they do not need to become the first-hop gateway for every tenant. Keeping roles clear reduces state and makes failure analysis more deterministic.

VTEP loopbacks connect the two layers

The VTEP source address is typically associated with a loopback that the underlay advertises. That loopback gives the overlay a stable tunnel endpoint independent of any one physical interface. When traffic is VXLAN-encapsulated, the outer IP header uses VTEP reachability that the underlay understands.

This is the key dependency: if the remote VTEP loopback is unreachable, changing EVPN route policy will not repair the data path. Underlay reachability, next hops, MTU, and ECMP should be validated first.

BGP EVPN distributes overlay information

MP-BGP EVPN gives the fabric a structured way to advertise endpoint and IP-prefix information. In an iBGP design, spine switches can act as route reflectors for EVPN routes while still forwarding data-plane traffic between leaves. The control plane tells VTEPs where remote reachability lives; VXLAN provides the encapsulation that carries traffic there.

Troubleshooting should distinguish “route not learned” from “route learned but tunnel transport broken.” That simple distinction avoids mixing BGP policy, endpoint learning, and physical reachability into one vague fabric problem.

Anycast gateways keep first-hop routing close to workloads

The virtualization context in data-center virtualization helps explain why workload mobility changed network design. A distributed anycast gateway lets multiple leaf switches present the same default-gateway identity for a tenant segment, so traffic can be routed at the local leaf instead of hairpinning to a centralized gateway.

This supports efficient east-west routing and reduces dependence on workload location. It also means gateway configuration and tenant VRF/VNI mapping must be consistent across the appropriate VTEPs.

MTU mistakes look like mysterious overlay failures

VXLAN adds encapsulation overhead. The underlay path must support a sufficiently large MTU so tenant traffic does not encounter fragmentation or drops unexpectedly. A ping with a small payload may succeed while larger application traffic fails, which can mislead an engineer into believing the overlay is healthy.

Validation should include path MTU and interface consistency across every leaf-spine hop. This is a good example of an overlay symptom caused by an underlay property.

Troubleshoot from physical reachability upward

The operational themes in CCNP Data Center preparation are strongest when applied as a layer-by-layer method. Verify interfaces and addressing, routing adjacencies, loopback reachability, ECMP paths, and MTU. Then verify NVE/VTEP state, BGP EVPN sessions and routes, VNI membership, endpoint learning, and finally tenant policy.

Packet captures or counters can confirm whether traffic is encapsulated, reaches the remote VTEP, and is decapsulated. The goal is to locate the first layer where expected state disappears.

The overlay-underlay model scales beyond memorizing commands

Commands and platform features change, but the separation remains durable. The routed fabric supplies resilient IP transport; the VXLAN EVPN overlay supplies logical segmentation, endpoint reachability, and tenant routing. Control-plane state tells the data plane how to use that transport.

Once those responsibilities are clear, design choices such as IGP versus eBGP underlay, route-reflector placement, anycast gateways, and border-leaf roles become easier to evaluate. The fabric stops looking like one giant protocol and becomes a set of layers with testable contracts.

Control-plane separation also helps during change windows. An engineer can validate underlay adjacencies and VTEP reachability before enabling tenant overlays, then verify EVPN peering and route exchange before attaching workloads. Staging checks in this order prevents several simultaneous unknowns from entering the fabric.

External connectivity introduces another boundary. Border leaves or border gateways connect the fabric to WAN, firewall, DCI, or other routing domains. Troubleshooting should still identify whether the problem is underlay transport, overlay reachability, or external route exchange. Border functions add policy but do not erase the layered model.

Automation benefits from the same abstraction. Underlay intent can define links, loopbacks, routing protocol, and MTU; overlay intent can define VRFs, networks, VNIs, gateways, and attachments. Separating those data models reduces accidental coupling and makes validation more targeted.

Addressing design deserves attention because loopbacks, point-to-point links, and optional unnumbered interfaces affect automation and troubleshooting. A structured plan makes it obvious which addresses represent device identity, VTEP reachability, and transit. When address roles are inconsistent, route tables become harder to interpret and scripted validation becomes more fragile.

Convergence should be tested during realistic failures. Pulling one leaf-spine link, restarting a routing process, or losing a spine can show whether the underlay reconverges within the application tolerance. Overlay sessions may remain established while data traffic briefly blackholes, or they may reset depending on the design. Operators should understand those behaviors before production incidents force the lesson.

ECMP troubleshooting requires checking more than route presence. A route can have several next hops while one physical path is dropping packets or using the wrong MTU. Interface counters, forwarding tables, and targeted traffic tests help verify that all equal-cost paths are healthy. Intermittent application failures are often explained by only some hashed flows landing on the bad path.

EVPN route types provide another diagnostic boundary. Different routes advertise MAC/IP reachability, inclusive multicast information, and IP prefixes. Engineers do not need to memorize every detail to benefit from the model: determine what reachability should be advertised, verify that the local VTEP originates it, confirm the control plane propagates it, and then confirm the remote VTEP installs usable forwarding state.

Operational documentation should preserve this layered troubleshooting order. A runbook that begins with random configuration commands encourages trial and error. A runbook that starts with underlay health, then overlay control plane, then tenant data plane gives less experienced engineers a repeatable method and reduces risky changes during incidents.

Anycast gateway behavior should be included in failure testing. If a leaf loses connectivity or a workload moves, the fabric should update endpoint reachability without forcing traffic through a remote gateway unnecessarily. Engineers can verify local gateway state, EVPN advertisements, and remote host routes during mobility or failure events. This turns the abstract promise of distributed routing into observable control-plane behavior.

Broadcast, unknown-unicast, and multicast handling is another overlay concern that should not be confused with ordinary underlay unicast. Depending on the design, the fabric may use multicast or ingress replication for BUM traffic. Troubleshooting should identify which replication method applies and verify the corresponding state rather than assuming every packet follows the same unicast path.

Finally, change control should validate both layers independently. An underlay routing change can affect every VTEP even when no tenant configuration changes, while an overlay policy change may affect only selected VRFs or VNIs. Separating blast radius in the change plan helps operators choose appropriate tests, rollback criteria, and maintenance windows.

Related Posts

• Start With Risk When Choosing Security Controls

• Why Azure VNets Fail: Address Spaces, Routes, and DNS

• NSGs, ASGs, and Azure Firewall: Put the Control in the Right Place

• Troubleshoot an Azure VM Before You Redeploy It

• Wireless Roaming, Channels, and the Physics of a Good WLAN

• Inside a Well-Designed Small Enterprise Network

• Prompt Management Becomes an Engineering Problem at Scale

• CI/CD for Prompts, Models, and AI Logic

• High Availability Is a System Property

• Multicast Without Mystery