← All writing
Fabric▶ Interactive

Where NCCL traffic actually goes: rail-optimized vs fat-tree

The same 32 GPUs, two ways to cable them, three kinds of traffic. Route them and count what has to cross the spine.

Two ways to cable a GPU cluster

Modern GPU servers have one high-speed NIC per GPU, so an 8-GPU node has 8 network ports. There are two common ways to connect them:

Which one wins depends entirely on who talks to whom. So route some traffic.

Topology explorer · 4 nodes × 8 GPUsline weight = flows on that link
Fabric Traffic
NIC ↔ leaf leaf ↔ spine traced flow
Network flows
–
Cross the spine
–
Switch hops / flow
–
average
Busiest link
–

What the explorer shows

Data-parallel rings stay on their rail

For each GPU, NCCL uses the NIC closest to it on PCIe, and it builds inter-node rings between same-index GPUs: GPU 3 on node 0 sends to GPU 3 on node 1. On a rail-optimized fabric every one of those flows is leaf-local, one switch hop and nothing on the spine. On the ToR fabric, the flows between nodes on different leaves must climb to a spine, and they share those uplinks with everything else. That is the case rail-optimized fabrics are built for, and why the spine layer can be lighter (or oversubscribed) in a rail design.

Tensor parallelism never touches the network

TP all-reduces happen every layer and are latency-critical, so frameworks keep TP groups inside a node, on NVLink. Switch to “TP in node” and the fabric goes quiet in both designs. Parallelism layout is a network design decision.

All-to-all is where rails hurt, and PXN is the fix

Mixture-of-experts layers send tokens from every GPU to every expert, so GPU i talks to GPU j ≠ i on other nodes. On rails, that means rail i → spine → rail j, and 7/8 of the flows land on the spine. NCCL 2.12 introduced PXN (PCI × NVLink): the source GPU first moves the data over NVLink to the GPU on the same node that sits on rail j, and that GPU sends it over its own NIC. Turn PXN on: spine flows go to zero, and the cost moves to NVLink, which has far more bandwidth than a NIC. It also lets NCCL aggregate messages headed for the same destination.

Why this matters beyond bandwidth

Keep going

Interactive
One dead link, 512 idle GPUs: ring all-reduce under failure
Interactive
The packet path: kernel stack vs XDP vs DPDK vs RDMA