Reliability at scale
goodput fault-tolerant training orchestrator
- Problem
- One bad GPU, NIC, cable or switch port stalls a collective, and every rank times out the same way.
- Why it’s hard
- Symptoms are global and causes are local. Stock tooling detects in minutes, attributes nothing, and restarts the whole job, sometimes onto the same broken hardware.
- Approach
- Inject faults with a ground-truth ledger, capture NCCL flight-recorder signatures, and correlate them with DCGM and NIC telemetry to name the culprit rank and component. Correct attribution makes small recoveries safe: re-init communicators in-process, replace one node, or restore from a peer’s in-memory checkpoint.
- Status
- M1 of 6 · in progress No results measured yet; every future number will link to a run.