Where NCCL traffic actually goes: rail-optimized vs fat-tree
The same 32 GPUs, two ways to cable them, three kinds of traffic. Route them and count what has to cross the spine.
Two ways to cable a GPU cluster
Modern GPU servers have one high-speed NIC per GPU, so an 8-GPU node has 8 network ports. There are two common ways to connect them:
- Top-of-rack (ToR) fat-tree. All 8 NICs of a node go to the same leaf switch, and leaves connect through spines. It is simple and rail-agnostic: traffic between nodes on different leaves goes up through a spine.
- Rail-optimized. NIC 0 of every node goes to leaf (rail) 0, NIC 1 to rail 1, and so on. NVIDIA’s NCCL 2.12 write-up describes it as maximizing all-reduce performance while minimizing interference between flows.
Which one wins depends entirely on who talks to whom. So route some traffic.
What the explorer shows
Data-parallel rings stay on their rail
For each GPU, NCCL uses the NIC closest to it on PCIe, and it builds inter-node rings between same-index GPUs: GPU 3 on node 0 sends to GPU 3 on node 1. On a rail-optimized fabric every one of those flows is leaf-local, one switch hop and nothing on the spine. On the ToR fabric, the flows between nodes on different leaves must climb to a spine, and they share those uplinks with everything else. That is the case rail-optimized fabrics are built for, and why the spine layer can be lighter (or oversubscribed) in a rail design.
Tensor parallelism never touches the network
TP all-reduces happen every layer and are latency-critical, so frameworks keep TP groups inside a node, on NVLink. Switch to “TP in node” and the fabric goes quiet in both designs. Parallelism layout is a network design decision.
All-to-all is where rails hurt, and PXN is the fix
Mixture-of-experts layers send tokens from every GPU to every expert, so GPU i talks to GPU j ≠ i on other nodes. On rails, that means rail i → spine → rail j, and 7/8 of the flows land on the spine. NCCL 2.12 introduced PXN (PCI × NVLink): the source GPU first moves the data over NVLink to the GPU on the same node that sits on rail j, and that GPU sends it over its own NIC. Turn PXN on: spine flows go to zero, and the cost moves to NVLink, which has far more bandwidth than a NIC. It also lets NCCL aggregate messages headed for the same destination.
Why this matters beyond bandwidth
- Blast radius. On rails, a bad rail switch or cable hits one GPU index on many nodes, the “rail-3 ranks are stuck” signature in my triage drill. On ToR, a bad leaf takes whole nodes.
- Congestion. Spine links are where flows from unrelated jobs meet. Keeping collectives leaf-local reduces how often pause back-pressure spreads between them.
- Attribution. Mapping “rank 11 is slow” to “node 1, NIC mlx5_3, rail switch 3, port 1” is exactly the topology join that nicprof and goodput depend on.