Naresh Kumar · GPU cluster networking & distributed-training infrastructure

A GPU cluster is only as fast as its slowest link and as reliable as its weakest. I make it lose less time to both.

~9 years below the framework layer: kernel networking, RDMA/RoCE, DPDK, SR-IOV, KVM and NCCL. Now applying it to the two problems that decide what a training or inference cluster actually delivers: reliability and efficiency at scale.

GPU cluster networking Distributed training NCCL RDMA / RoCE v2 Disaggregated inference DPDK SR-IOV / KVM eBPF / XDP Linux kernel
Fig. 0 · Ring all-reduce under failure8 GPUs
Ring all-reduce losing a link and recovering Eight GPUs pass gradient chunks around a ring. The link between GPU 2 and GPU 3 fails, the stall spreads around the ring, every rank times out, and the job restores from a checkpoint onto a spare GPU. A goodput timeline shows productive time in blue, stalled time hatched and restart time in red.

Every second the bar isn’t blue is lost goodput. Break the ring yourself →

Repo · reliability↗
goodput: fault attribution & recovery for multi-node training
Design, interfaces, test harness public. M1 of 6.
Repo · runnable↗
nicprof: which NIC counter explains which millisecond of step time
nicprof demo · 69 tests
Repo · efficiency↗
kvwire: RDMA KV-cache transport for disaggregated inference
vLLM KV-connector plug-in. M1 in progress.
Hugging Face↗
6 root-cause & failure-prediction models
NCCL, GPU, network, Linux and Kubernetes logs

Work

Two problems I work on

AI lab infrastructure splits along two lines: keep the job running, and make each GPU-hour count. Kernel and fabric depth is the foundation for both.

Reliability at scale

goodput fault-tolerant training orchestrator

Problem
One bad GPU, NIC, cable or switch port stalls a collective, and every rank times out the same way.
Why it’s hard
Symptoms are global and causes are local. Stock tooling detects in minutes, attributes nothing, and restarts the whole job, sometimes onto the same broken hardware.
Approach
Inject faults with a ground-truth ledger, capture NCCL flight-recorder signatures, and correlate them with DCGM and NIC telemetry to name the culprit rank and component. Correct attribution makes small recoveries safe: re-init communicators in-process, replace one node, or restore from a peer’s in-memory checkpoint.
Status
M1 of 6 · in progress  No results measured yet; every future number will link to a run.

Efficiency at scale

kvwire disaggregated inference engine

Problem
Prefill is compute-bound, decode is memory-bandwidth-bound. Colocated, they interfere: prefills spike TPOT, decodes spike TTFT.
Why it’s hard
Disaggregation trades interference for data movement. The KV cache has to cross the network before decode can start, and has to stay correct when the network fails mid-transfer.
Approach
A GPUDirect RDMA transport, a Triton KV re-layout kernel (kept only if the data justifies it), and a failure-aware coordinator. It plugs into vLLM through the KV connector interface and replaces the parts that move bytes and decide what to do when moving bytes goes wrong.
Status
M1 · in progress  Baselines and hardware ceilings first; results cells stay “—” until measured.

Writing

Interactive notes & posts

Interactive posts run in the browser: break a ring, size a KV transfer, freeze a fabric. Every model is labeled with its assumptions. Long-form essays live on Substack.

Upstream & evidence

Things you can run or read

Each entry says what problem it addresses and links to the source.

NIC counter → step-time attribution nicprof

The gap: a 512-GPU job slows down; the network team has millions of counter series, the ML team has a step-time graph, and nobody can say which counter explains which millisecond.

What it does: aligns RDMA/ethtool/SR-IOV counters to each rank’s step windows, fits a non-negative model of excess step time, and traces PFC pause back to its origin so a victim port isn’t blamed as the cause.

Runnable demoSource ↗

Kernel performance toolkit kptk

The gap: “CPU 80%” and “THP on” don’t explain why a workload is slow.

What it does: explains slowness from kernel evidence (run-queue delay, allocation vs access NUMA locality, THP fallback and compaction stalls) and calls perf_event_open(2) directly with multiplexing correction. Standard library only, 43 tests.

Runnable demoSource ↗

Upstream: Linux memory manager UAPI headers PR #10

What changed: fixed typos in the UAPI headers of an open-source Linux memory manager project. Small, and merged.

MergedPR ↗

Foundations

Skill → the artifact that shows it

No self-ratings. Each area links to the repo or write-up behind it, with an honest stage.

AreaStageEvidence
Collectives & NCCLdesign + harness goodput ↗ distributed-training-framework-nccl ↗ ring all-reduce post rail topology post
RDMA / RoCE / PFCrunnable nicprof ↗ kvwire ↗ PFC deadlock post packet path post
Linux kernel performancerunnable kptk ↗
Inference servingearly kvwire ↗ high-performance-llm-inference-engine ↗ KV-cache post
GPU kernels (CUDA / Triton)early custom-cuda-fused-attention-triton ↗
eBPF / XDP / tcearly kernel-level-ai-traffic-shaper ↗ packet path post
DPDKwrite-up DPDK essay ↗ packet path post
Virtualization (SR-IOV / KVM)write-up Virtualization essay ↗
ML for infra opspublished 6 HF models ↗ 6 HF Spaces ↗

Profiles

Proof of work & where to find me

Technical proof of work

Social & writing

Open to roles

Building training or inference clusters? Let’s talk.

I’m looking for GPU cluster networking, distributed-training infrastructure and inference-platform roles at AI labs and infra teams.

USACanadaEurope