NVIDIA NCA-AIIO: GPU Cluster Topology for AI Workloads
GPU topology determines which paths data can take between accelerators, CPUs, memory, network interfaces, and storage. On a single node, PCIe layout and high-speed GPU interconnects shape peer-to-peer communication. Across nodes, the network fabric determines whether distributed training and inference can exchange data fast enough to keep GPUs busy.
For candidates following NCA-AIIO, topology is a foundational operations concept because accelerated computing performance depends on more than the model of GPU installed. Inside AI infrastructure, topology is the map that explains why two clusters with the same accelerator count can behave very differently under the same distributed workload.
Topology decisions should follow the communication pattern. Data-parallel training, tensor or pipeline parallelism, distributed inference, storage-heavy preprocessing, and independent batch jobs place different demands on scale-up and scale-out links. The right architecture minimizes expensive communication for the workload the platform is meant to run.
Distinguish scale-up from scale-out communication
Scale-up communication happens among accelerators within a tightly coupled system, while scale-out communication crosses node boundaries. High-bandwidth GPU interconnects can make intra-node exchange much faster than ordinary host paths, so placement decisions that keep heavily communicating ranks within the same fast domain can materially improve performance.
Scale-out communication depends on network interfaces, switches, routing, congestion behavior, and the collective library. As node count increases, the fabric must carry more simultaneous exchanges. The architecture should identify where bandwidth is dedicated, where it is shared, and what oversubscription the workload can tolerate.
Infrastructure bottlenecks often become visible only when an application moves from one node to many. A topology review helps explain whether the slowdown comes from communication distance, limited network bandwidth, a shared uplink, or a software path that is not using the intended hardware.
Map PCIe, NUMA, and NIC locality
Inside a server, not every GPU has the same relationship to every CPU socket or network adapter. Traffic that crosses a CPU interconnect or remote PCIe path can add latency and reduce available bandwidth. Dense nodes should therefore be documented with GPU-to-NIC and GPU-to-CPU locality rather than represented as a box containing a number of identical devices.
Schedulers and distributed frameworks can use topology information only if the platform exposes it consistently. Node labels, device discovery, and workload placement policies should reflect the actual hardware. If administrators replace cards or change firmware and the topology description becomes stale, placement decisions can quietly move traffic onto slower paths.
NUMA awareness matters for host memory as well. Data loaders and communication processes should avoid unnecessary remote-memory access when the workload is sensitive to latency. The goal is not to micro-optimize every process manually but to make the normal placement policy align with the physical layout.
Design east-west networking for collective traffic
Distributed AI sends large volumes of east-west traffic between compute nodes. Collective operations can create synchronized bursts and many-to-many patterns that are very different from conventional client-server application traffic. A network designed mainly for north-south access may have enough aggregate port bandwidth and still perform poorly under these patterns.
NVIDIA reference architectures emphasize balanced east-west bandwidth and, at larger scales, separate traffic classes or fabrics so compute communication does not compete unpredictably with storage, management, and external client traffic. The exact thresholds depend on the architecture, but the principle is stable: high-value accelerator traffic should have a path designed for its concurrency.
Network monitoring should measure congestion and path utilization during real collectives. A link that averages 30 percent utilization over a minute can still experience microbursts that stall synchronized workers. Monitoring intervals and metrics should be chosen to reveal the workload behavior the team is trying to protect.
Keep storage paths visible in the topology
AI clusters often focus on GPU-to-GPU communication and treat storage as an external box. That is risky because training, checkpointing, model loading, and data preprocessing can generate intense north-south or storage-network traffic. Storage endpoints and their network paths should appear on the same topology diagram as compute fabrics.
Separate storage and compute traffic when contention justifies it, or at minimum understand how they share links. A checkpoint burst that collides with an all-reduce phase can look like random training instability if the architecture hides the shared switch or uplink.
Stateful workloads also matter in Kubernetes. Pods can be rescheduled freely, but large datasets and model artifacts have placement and recovery costs. Topology-aware storage choices can reduce unnecessary data movement and make failure recovery more predictable.
Use failure domains when choosing placement
Topology is also an availability model. Racks, power domains, switches, network planes, and nodes fail independently or together. A job that spans more failure domains may tolerate a single component failure differently from one concentrated on a rack, while a highly coupled training run may prefer locality for performance. The scheduler and platform should know which trade-off is being made.
For serving workloads, replicas can be spread across failure domains so one rack or switch does not remove the entire service. For large training jobs, checkpoint strategy may be more important because losing one rank can restart a coordinated job. Topology and recovery design should therefore be evaluated together.
Dual-plane or redundant network designs improve resilience only if software can use the alternate path. Test failover under workload and measure the performance of the degraded state. A secondary path that technically reconnects but cannot sustain acceptable communication may meet a wiring requirement without meeting the service objective.
Topology-aware scheduling needs accurate resource descriptions
GPU scheduling becomes more effective when resource requests describe what the job actually needs: accelerator type, memory capacity, count, sharing mode, and sometimes topology affinity. If every request is simply “one GPU,” the scheduler has little information with which to preserve scarce high-bandwidth groups for jobs that need them.
Multi-Instance GPU and time-slicing add another topology layer. MIG partitions supported GPUs into isolated instances, while time-slicing shares a physical GPU without the same memory and fault isolation. The scheduling system must expose these resources clearly so users understand whether they are requesting dedicated or shared capacity.
Kubernetes RBAC remains relevant because the objects that control node labels, device plugins, and operator configuration are high-impact infrastructure controls. Topology-aware scheduling is only reliable when those descriptions are governed and changes are auditable.
Validate topology with workload telemetry
A topology diagram is a hypothesis about performance. Validate it with device and network telemetry while representative jobs run. Measure GPU communication activity, PCIe or NVLink traffic, network throughput, collective time, CPU use, and storage throughput. The objective is to confirm that traffic uses the intended fast paths.
GPU monitoring should be correlated with topology because utilization alone cannot show why communication stalls. DCGM profiling metrics can expose activity across compute, memory, PCIe, and NVLink, giving operators evidence about whether the workload is compute-bound or communication-bound.
Repeat validation after hardware replacement, firmware changes, driver upgrades, and network redesign. Topology is physical, but the effective data path depends on software configuration. A change that leaves the cabling untouched can still alter which route the workload takes.
Document topology as an operational asset
Topology documentation should be generated from the live environment where possible. Static diagrams become stale after a NIC replacement, rack expansion, cable move, or firmware change. Inventory systems can record GPU, NIC, switch, rack, and node relationships, while validation jobs can confirm that peer-to-peer and network paths perform within the expected range.
Give operators a way to move from a slow job to the physical path quickly. They should be able to identify the GPUs assigned to the job, the node and rack containing them, the relevant NICs, the switches carrying east-west traffic, and the storage path used for checkpoints. This reduces the time spent correlating several team-specific inventories during an incident.
Topology knowledge also improves maintenance planning. Draining one rack, switch plane, or group of nodes can remove a disproportionate amount of high-bandwidth capacity even if the raw GPU count remaining looks adequate. Maintenance reviews should therefore consider the shapes of topology that remain schedulable, not only total devices online.
Topology standards should be reviewed whenever the workload portfolio changes. A fabric tuned for data-parallel training may not be the best fit for large distributed inference or storage-intensive retrieval workloads. Rather than redesigning for every model, define a small set of supported topology patterns and measure which workload classes fit each one. This gives schedulers and capacity planners a practical vocabulary for placement and growth.
Capacity models should record topology scarcity explicitly. If only a small subset of nodes can provide the fastest GPU-to-GPU or GPU-to-NIC path, those nodes are a specialized resource even when the hardware model matches the rest of the fleet. Tracking demand for that topology helps prevent ordinary jobs from consuming the exact placement needed by tightly coupled workloads.
GPU cluster topology is the structure that turns a collection of accelerators into a usable compute system. It defines locality, communication distance, failure domains, storage paths, and the network capacity available to distributed workloads.
The strongest designs make those relationships visible to both schedulers and operators. When placement and telemetry align with the physical topology, teams can scale workloads with fewer surprises and can explain performance changes with evidence rather than guesswork.