← All writing
Reliability▶ Interactive drill

Reading the crime scene: triaging an NCCL timeout

32 ranks, one stalled all-reduce, 32 identical error messages. Pull evidence one source at a time and name the culprit before you restart anything.

The setup

A 4-node job, 8 GPUs per node, 32 ranks, data-parallel with NCCL. GPU i on every node uses NIC mlx5_i (a rail-optimized layout), and rank = node × 8 + gpu. Something went wrong mid-step. Ten minutes later, every rank’s watchdog fired.

Pick an incident, look at the cluster, and open evidence sources in whatever order you think is fastest. Fewer sources used is better: in a real outage every query costs minutes.

Incident drill 0 of 4 sources used

watchdog timeout flight recorder: behind or stuck on network never enqueued this op
Your call

The method, without the game

  1. Ignore the loudest symptom. Identical timeouts on every rank confirm that the collective stalled. They don’t say where.
  2. Find the break in collective state. Compare the last enqueued, started and completed sequence numbers across ranks. A rank that never enqueued the op points at the host or application side. A rank that’s behind, or a group of ranks sharing a rail, points at a device or the fabric.
  3. Correlate lower layers in the same time window. GPU Xid events and DCGM health for the suspect GPUs; NIC link-down, symbol-error, discard and pause counters for the suspect rails.
  4. Separate victims from causes. Back-pressure spreads, so the noisiest counters can sit on innocent ports (see PFC).
  5. Choose the smallest safe action for what you found, then resume.

None of this is exotic, but done by hand it takes a person, many dashboards and a lot of minutes per incident. Automating it is the core of goodput: capture flight-recorder signatures, join them with DCGM and NIC telemetry on a topology map, and output a culprit with its evidence. nicprof already does the NIC-counter half for slowdowns, and the NCCL-RCA model on Hugging Face is my experiment in learning the log-reading half.

Keep going

Next · interactive
How often should a 16K-GPU job checkpoint?
Repo
goodput: fault attribution & recovery ↗