← All writing
Foundations▶ Interactive

The packet path: kernel stack vs XDP vs DPDK vs RDMA

Four ways for bytes to get from a NIC to the code that wants them. Each one moves work away from the general-purpose kernel, and each one gives something up to do it.

The budget

Start with arithmetic. On Ethernet, every frame also costs 20 bytes on the wire (preamble, start delimiter and inter-frame gap). So a packet of S bytes occupies (S + 20) × 8 bit-times, and at line rate the next one arrives right behind it. Divide by your clock and you get the CPU cycles one core can spend on each packet before it falls behind.

Line-rate budget per packetpure arithmetic
Packets / s at line rate
–
Time per packet
–
Cycles per packet
–

A few dozen cycles is less than one cache miss to DRAM. The only ways through are to do less per packet, touch fewer cache lines, batch work, spread it across many cores, or hand it to the NIC. The four paths below are different mixes of those moves.

Four paths, one figure

Pick a path and step through it. The bands show where each step runs: NIC hardware, the kernel, or user space. Dashed boxes are setup the kernel does once, not per packet.

One packet, four datapathsStep through, or play

What each path gives up

Kernel stackXDP / AF_XDPDPDKRDMA
Per-packet work runs inkernel (softirq)driver hook, then kernel or useruser-space poll loopNIC hardware
CPU copies to the app1 (socket → user buffer)0 with AF_XDP zero-copy0 (app reads the mbuf)0, and none on the remote CPU
Interrupts on the data pathyes, mitigated by NAPIsame as kernel (NAPI)none (busy polling)none needed (poll the CQ)
Transport (TCP, retransmit)kernelkernel on XDP_PASS, else yoursyours (or a user-space stack)NIC: reliable connection in hardware
tcpdump, iptables, ss…all workpartly (before the stack)no: the kernel never sees the packetsno: needs NIC-side counters
What you give upper-packet efficiencyverifier limits, per-packet programsdedicated cores at 100%, your own stackverbs API, memory registration, a fabric that tolerates RDMA

Where this shows up in an AI cluster

Keep going

Next · interactive
Where NCCL traffic actually goes: rail-optimized vs fat-tree
Interactive
PFC: the lossless network that can freeze itself