← All writing
Fabric▶ Interactive

PFC: the lossless network that can freeze itself

RoCE wants a network that never drops a packet. Priority Flow Control delivers that by letting a switch tell its neighbour to stop. Arrange the “stop”s in a cycle and nobody ever starts again.

Why RoCE wants a lossless fabric

RDMA NICs implement the transport in hardware, and RoCE v2’s classic loss recovery is go-back-N: one drop can trigger a retransmission of everything after it. In a training cluster, collectives move gradient buffers that can be gigabytes, so a little loss means a large throughput hit. Operators therefore run RDMA traffic in a lossless priority class using PFC (IEEE 802.1Qbb).

The mechanism is simple. When a switch’s ingress buffer for a priority fills past XOFF, it sends a PAUSE frame upstream for that priority. When the buffer drains below XON, it sends a resume. Packets wait instead of being dropped.

How waiting turns into deadlock

A paused upstream switch keeps receiving traffic, so its own buffers fill and it pauses its upstream. Back-pressure spreads hop by hop. Usually it ends at a host NIC, which simply sends less. But if the waits form a cycle (switch A waiting on B, B on C, C on D, D on A), every buffer is full of packets that can only leave through a paused link. Nothing drains, so nothing resumes.

Clos fabrics with up-down routing don’t have such cycles in steady state. They appear through link failures and reroutes (packets bouncing leaf → spine → leaf → spine), through misconfiguration, or through flooding. The toy below forces the cycle with four flows that each cross two hops around a ring of switches.

Toy lossless fabric Run, then hit “Burst”
transit packet (needs next link) local packet (exits here) link sending link paused
Fabric state
–
Delivered / tick
–
8-tick average
PAUSE frames sent
–
Paused links

Press Run. At moderate load, queues rise and fall and pauses come and go. Press Burst to make every host dump a full send queue at once.

What to try

  1. Run, then Burst. Within a few ticks every link is paused and throughput drops to zero. Now turn Hosts sending off. The deadlock stays: the packets that would drain are behind packets that can’t move.
  2. Switch to 3 flows. Removing one flow breaks the cycle. One switch now holds only local traffic, which always drains, so the chain of pauses always has an end. Pauses still happen. Deadlock can’t.
  3. Back to 4 flows, turn ECN on, Burst. Switches mark packets once a queue passes a lower threshold (Kmin), and senders cut their rate before buffers reach XOFF. PFC becomes a rare safety net instead of the main flow control.

How production fabrics deal with it

The debugging trap: victims look like culprits

Look at the pause counters during a storm. Pauses propagate upstream, so the switches and NICs furthest from the cause can show the highest pause counts. A dashboard sorted by “most pause frames sent” points straight at a victim. Attribution has to follow the dependency back to the port where the backlog started.

That is one of the things nicprof does. It ties NIC counter evidence to excess training step time, and it traces pause propagation back to its origin so the loudest port isn’t automatically blamed. It’s also why my disaggregated inference work treats congestion as part of the transport design, not an afterthought.

Keep going

Interactive
One dead link, 512 idle GPUs: ring all-reduce under failure
Repo
nicprof: NIC counters → step time ↗