PFC: the lossless network that can freeze itself
RoCE wants a network that never drops a packet. Priority Flow Control delivers that by letting a switch tell its neighbour to stop. Arrange the “stop”s in a cycle and nobody ever starts again.
Why RoCE wants a lossless fabric
RDMA NICs implement the transport in hardware, and RoCE v2’s classic loss recovery is go-back-N: one drop can trigger a retransmission of everything after it. In a training cluster, collectives move gradient buffers that can be gigabytes, so a little loss means a large throughput hit. Operators therefore run RDMA traffic in a lossless priority class using PFC (IEEE 802.1Qbb).
The mechanism is simple. When a switch’s ingress buffer for a priority fills past XOFF, it sends a PAUSE frame upstream for that priority. When the buffer drains below XON, it sends a resume. Packets wait instead of being dropped.
How waiting turns into deadlock
A paused upstream switch keeps receiving traffic, so its own buffers fill and it pauses its upstream. Back-pressure spreads hop by hop. Usually it ends at a host NIC, which simply sends less. But if the waits form a cycle (switch A waiting on B, B on C, C on D, D on A), every buffer is full of packets that can only leave through a paused link. Nothing drains, so nothing resumes.
Clos fabrics with up-down routing don’t have such cycles in steady state. They appear through link failures and reroutes (packets bouncing leaf → spine → leaf → spine), through misconfiguration, or through flooding. The toy below forces the cycle with four flows that each cross two hops around a ring of switches.
Press Run. At moderate load, queues rise and fall and pauses come and go. Press Burst to make every host dump a full send queue at once.
What to try
- Run, then Burst. Within a few ticks every link is paused and throughput drops to zero. Now turn Hosts sending off. The deadlock stays: the packets that would drain are behind packets that can’t move.
- Switch to 3 flows. Removing one flow breaks the cycle. One switch now holds only local traffic, which always drains, so the chain of pauses always has an end. Pauses still happen. Deadlock can’t.
- Back to 4 flows, turn ECN on, Burst. Switches mark packets once a queue passes a lower threshold (Kmin), and senders cut their rate before buffers reach XOFF. PFC becomes a rare safety net instead of the main flow control.
How production fabrics deal with it
- Remove cycles: up-down routing, and care with reroutes after link failures so packets never bounce back up.
- Keep PFC as a backstop: ECN with DCQCN (or delay-based schemes) keeps queues short, so PFC rarely fires. This is the same reason DCQCN thresholds sit well below XOFF.
- Watch for storms: switch and NIC PFC watchdogs detect a queue that stays paused too long and drop or disable lossless mode on it, trading a few drops for liveness.
- Make the NIC more loss-tolerant: newer RDMA NICs support selective retransmission, which reduces the need for a strictly lossless fabric.
The debugging trap: victims look like culprits
Look at the pause counters during a storm. Pauses propagate upstream, so the switches and NICs furthest from the cause can show the highest pause counts. A dashboard sorted by “most pause frames sent” points straight at a victim. Attribution has to follow the dependency back to the port where the backlog started.
That is one of the things nicprof does. It ties NIC counter evidence to excess training step time, and it traces pause propagation back to its origin so the loudest port isn’t automatically blamed. It’s also why my disaggregated inference work treats congestion as part of the transport design, not an afterthought.