CareerPlanSign in

Software Engineer - GPU Networking & Distributed Systems

San Francisco💼 Full-time🗓 2026-02-23 → 2026-09-25

Core

Architecting the software fabric for distributed, heterogeneous AI hardware to unify thousands of GPUs and optimize communication for LLM inference.

Role type

Senior IC GPU Networking & Distributed Systems Engineer

Builds

Software primitives for Disaggregated Serving, Wide Expert Parallelism (WideEP), and low-latency distributed inference stacks.

Domain

AI Infrastructure / High-Performance Computing / GPU Networking

Deliverable

production ML models

Required skills

High-performance networking protocols (InfiniBand, RoCE v2), C++, Python, NVIDIA architecture memory hierarchy, custom kernel development, distributed system debugging

Preferred skills

NCCL, NVSHMEM, UCX, Rust, GPUDirect Storage, TensorRT-LLM, vLLM, SGLang

Technologies

InfiniBand, RoCE, NVLink, NCCL, NVSHMEM, TensorRT-LLM, Kubernetes

Responsibilities

Integrate RDMA/RoCE/InfiniBand into the inference stack; implement networking layers for Disaggregated KV Cache Offload and WideEP; optimize checkpointing for sub-10-second LLM startup; characterize and validate networking performance on H100/B200/NVL72 clusters; design observability tools for packet flow and congestion; write custom communication kernels to overlap compute and data transfer.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.