← All writing
Reliability▶ Interactive

How often should a 16K-GPU job checkpoint?

Checkpoint too often and you pay for writes. Too rarely and every failure throws away hours. A 1974 formula gives the sweet spot, and it also shows which engineering lever is worth the most.

Failures scale with the job

If each GPU (with its share of host, NIC, cables and switch ports) fails independently with mean time between failures m, a job spanning N of them is interrupted every M = m / N on average. Any one failure stops a synchronous training job.

For a real reference point, Meta’s Llama 3 paper reports 419 unexpected interruptions over a 54-day snapshot of pretraining on 16,384 H100s, most attributed to hardware. That is one interruption roughly every 3 hours, which works out to an effective per-GPU MTBF of about 5.8 years. The calculator’s default preset uses that rate.

The trade-off

Each checkpoint costs δ of stalled training. Each failure costs a restart R (detect, reschedule, init, load) plus, on average, half an interval of lost work. Per unit time the waste is roughly:

waste(τ) ≈ δ/τ + (τ/2 + R) / M → τ* ≈ √(2·δ·M)

This is Young’s first-order result; Daly refined it for cases where τ isn’t small relative to M. Goodput is 1 − waste.

Goodput vs checkpoint interval Young/Daly first-order model
Job MTBF
–
Optimal interval
–
Best goodput
–
At your interval
–
goodput, current settings same, with restart cost ÷ 5 optimum your interval

Which lever is worth the most?

Improvement (re-optimized interval)Best goodputGainGPU-hours / day back

Reading the sensitivity table

The ranking depends on where you start, and that is the point. With the defaults (2-minute blocking checkpoints at 16K GPUs), making checkpoints asynchronous or in-memory is the biggest lever. Now set the checkpoint stall to 10 seconds, which is what async and in-memory checkpointing aim for, and look again. With cheap checkpoints, the optimum moves to a few minutes and the remaining waste is mostly the restart cost paid on every failure. Cutting it 5× becomes the top lever, ahead of halving the failure rate.

Restart cost is mostly not hardware:

The same levers shrink lost work: in-memory or peer checkpoints make δ small enough to checkpoint every few minutes.

Where this goes

goodput targets the restart-cost term directly: fast, correct attribution to the culprit rank and component, then the smallest safe recovery (in-process communicator re-init, single-node replacement, peer in-memory restore). Its thesis is literally loss per fault ≈ MTTD + MTTR + lost work + false-positive cost. This page is the back-of-envelope version. The repo will measure it.

Keep going

Next · interactive
Reading the crime scene: triaging an NCCL timeout
Repo
goodput: fault attribution & recovery ↗