Network Engineer, Supercomputing
Core
Own the lowest layers of the network stack for large-scale AI training and inference, ensuring interconnect reliability across GPU fabrics.
Role type
Senior IC network engineer (supercomputing/HPC)
Builds
Production GPU fabrics (RDMA/RoCE, NVLink/NVSwitch) for AI training clusters
Domain
AI infrastructure / High-performance computing
Deliverable
production ML models
Required skills
Linux host-level debugging, Python or Rust, Kubernetes or Slurm, RDMA/RoCEv2, NVLink/NVSwitch, distributed systems
Preferred skills
Cloud network primitives, CUDA/NCCL, performance profiling, statistical rigor in reliability reasoning, tooling development
Technologies
RDMA, RoCEv2, NVLink, NVSwitch, NCCL, Kubernetes, Slurm, Python, Rust, Linux
Responsibilities
Validate GPU network fabric design, debug collective failures and congestion control, own NVLink/NVSwitch interconnect health, build network instrumentation dashboards, triage cross-cloud fabric issues, drive escalations with cloud providers
Seniority
Senior, hands-on IC