The packet path: kernel stack vs XDP vs DPDK vs RDMA
Four ways for bytes to get from a NIC to the code that wants them. Each one moves work away from the general-purpose kernel, and each one gives something up to do it.
The budget
Start with arithmetic. On Ethernet, every frame also costs 20 bytes on the wire (preamble, start delimiter and inter-frame gap). So a packet of S bytes occupies (S + 20) × 8 bit-times, and at line rate the next one arrives right behind it. Divide by your clock and you get the CPU cycles one core can spend on each packet before it falls behind.
A few dozen cycles is less than one cache miss to DRAM. The only ways through are to do less per packet, touch fewer cache lines, batch work, spread it across many cores, or hand it to the NIC. The four paths below are different mixes of those moves.
Four paths, one figure
Pick a path and step through it. The bands show where each step runs: NIC hardware, the kernel, or user space. Dashed boxes are setup the kernel does once, not per packet.
What each path gives up
| Kernel stack | XDP / AF_XDP | DPDK | RDMA | |
|---|---|---|---|---|
| Per-packet work runs in | kernel (softirq) | driver hook, then kernel or user | user-space poll loop | NIC hardware |
| CPU copies to the app | 1 (socket → user buffer) | 0 with AF_XDP zero-copy | 0 (app reads the mbuf) | 0, and none on the remote CPU |
| Interrupts on the data path | yes, mitigated by NAPI | same as kernel (NAPI) | none (busy polling) | none needed (poll the CQ) |
| Transport (TCP, retransmit) | kernel | kernel on XDP_PASS, else yours | yours (or a user-space stack) | NIC: reliable connection in hardware |
| tcpdump, iptables, ss… | all work | partly (before the stack) | no: the kernel never sees the packets | no: needs NIC-side counters |
| What you give up | per-packet efficiency | verifier limits, per-packet programs | dedicated cores at 100%, your own stack | verbs API, memory registration, a fabric that tolerates RDMA |
Where this shows up in an AI cluster
- Collectives run on RDMA. NCCL’s network transport uses InfiniBand verbs (over InfiniBand or RoCE), and with GPUDirect RDMA the NIC reads and writes GPU memory directly over PCIe, with no bounce through host memory. Toggle GPUDirect in the figure to see the extra hop it removes. It’s also the path my kvwire transport uses to move KV cache between prefill and decode GPUs.
- RDMA needs a fabric that cooperates. Hardware transport assumes little or no loss, which is why RoCE fabrics run PFC and ECN, and why NIC counters, not tcpdump, are how you debug them. That’s the gap nicprof works in.
- The kernel stack still carries everything else: control planes, checkpoint and dataset traffic over TCP, and inference front-ends. Its per-packet cost is why tuning (RSS, GRO, busy polling, IRQ affinity) matters at high rates. kptk is my toolkit for the CPU side of that.
- eBPF is the programmable layer: XDP for early drops and steering, tc/eBPF for shaping and policy. My traffic-shaper repo is an early experiment there.
- DPDK wins where packet rate dominates and you can afford dedicated cores, such as virtual switches, load balancers and packet processing appliances. I wrote about it on Substack.