How often should a 16K-GPU job checkpoint?
Checkpoint too often and you pay for writes. Too rarely and every failure throws away hours. A 1974 formula gives the sweet spot, and it also shows which engineering lever is worth the most.
Failures scale with the job
If each GPU (with its share of host, NIC, cables and switch ports) fails independently with mean time between failures m, a job spanning N of them is interrupted every M = m / N on average. Any one failure stops a synchronous training job.
For a real reference point, Meta’s Llama 3 paper reports 419 unexpected interruptions over a 54-day snapshot of pretraining on 16,384 H100s, most attributed to hardware. That is one interruption roughly every 3 hours, which works out to an effective per-GPU MTBF of about 5.8 years. The calculator’s default preset uses that rate.
The trade-off
Each checkpoint costs δ of stalled training. Each failure costs a restart R (detect, reschedule, init, load) plus, on average, half an interval of lost work. Per unit time the waste is roughly:
This is Young’s first-order result; Daly refined it for cases where τ isn’t small relative to M. Goodput is 1 − waste.
Reading the sensitivity table
The ranking depends on where you start, and that is the point. With the defaults (2-minute blocking checkpoints at 16K GPUs), making checkpoints asynchronous or in-memory is the biggest lever. Now set the checkpoint stall to 10 seconds, which is what async and in-memory checkpointing aim for, and look again. With cheap checkpoints, the optimum moves to a few minutes and the remaining waste is mostly the restart cost paid on every failure. Cutting it 5× becomes the top lever, ahead of halving the failure rate.
Restart cost is mostly not hardware:
- Detection: waiting for a collective timeout (10 minutes by default in PyTorch) before anything happens.
- Attribution: working out which node to drain. Get it wrong and the restarted job lands on the same bad link (see why every rank reports the same error).
- Blast radius: tearing down every process, reinitializing CUDA and NCCL on every rank, and reloading from remote storage, when one node failed.
The same levers shrink lost work: in-memory or peer checkpoints make δ small enough to checkpoint every few minutes.
Where this goes
goodput targets the restart-cost term directly: fast, correct attribution to the culprit rank and component, then the smallest safe recovery (in-process communicator re-init, single-node replacement, peer in-memory restore). Its thesis is literally loss per fault ≈ MTTD + MTTR + lost work + false-positive cost. This page is the back-of-envelope version. The repo will measure it.