Reading the crime scene: triaging an NCCL timeout
32 ranks, one stalled all-reduce, 32 identical error messages. Pull evidence one source at a time and name the culprit before you restart anything.
The setup
A 4-node job, 8 GPUs per node, 32 ranks, data-parallel with NCCL. GPU i on every node uses NIC mlx5_i (a rail-optimized layout), and rank = node × 8 + gpu. Something went wrong mid-step. Ten minutes later, every rank’s watchdog fired.
Pick an incident, look at the cluster, and open evidence sources in whatever order you think is fastest. Fewer sources used is better: in a real outage every query costs minutes.
The method, without the game
- Ignore the loudest symptom. Identical timeouts on every rank confirm that the collective stalled. They don’t say where.
- Find the break in collective state. Compare the last enqueued, started and completed sequence numbers across ranks. A rank that never enqueued the op points at the host or application side. A rank that’s behind, or a group of ranks sharing a rail, points at a device or the fabric.
- Correlate lower layers in the same time window. GPU Xid events and DCGM health for the suspect GPUs; NIC link-down, symbol-error, discard and pause counters for the suspect rails.
- Separate victims from causes. Back-pressure spreads, so the noisiest counters can sit on innocent ports (see PFC).
- Choose the smallest safe action for what you found, then resume.
None of this is exotic, but done by hand it takes a person, many dashboards and a lot of minutes per incident. Automating it is the core of goodput: capture flight-recorder signatures, join them with DCGM and NIC telemetry on a topology map, and output a culprit with its evidence. nicprof already does the NIC-counter half for slowdowns, and the NCCL-RCA model on Hugging Face is my experiment in learning the log-reading half.