Software Engineer - GPU Networking & Distributed Systems
Core
Architecting the software fabric for distributed, heterogeneous AI hardware to unify thousands of GPUs and optimize communication for LLM inference.
Role type
Senior IC GPU Networking & Distributed Systems Engineer
Builds
Software primitives for Disaggregated Serving, Wide Expert Parallelism (WideEP), and low-latency distributed inference stacks.
Domain
AI Infrastructure / High-Performance Computing / GPU Networking
Deliverable
production ML models
Required skills
High-performance networking protocols (InfiniBand, RoCE v2), C++, Python, NVIDIA architecture memory hierarchy, custom kernel development, distributed system debugging
Preferred skills
NCCL, NVSHMEM, UCX, Rust, GPUDirect Storage, TensorRT-LLM, vLLM, SGLang
Technologies
InfiniBand, RoCE, NVLink, NCCL, NVSHMEM, TensorRT-LLM, Kubernetes
Responsibilities
Integrate RDMA/RoCE/InfiniBand into the inference stack; implement networking layers for Disaggregated KV Cache Offload and WideEP; optimize checkpointing for sub-10-second LLM startup; characterize and validate networking performance on H100/B200/NVL72 clusters; design observability tools for packet flow and congestion; write custom communication kernels to overlap compute and data transfer.
Seniority
Senior, hands-on IC
