GPU Cluster Engineer, Networking
Core
Design, build, and operate the complete networking stack for large-scale GPU clusters powering multimodal AI models, including RDMA fabrics, data center perimeters, and cross-site connectivity.
Role type
Senior Network Engineer (AI Infrastructure)
Builds
High-performance GPU cluster fabrics (InfiniBand/RoCE), data center networks, and hybrid cloud interconnects for frontier AI workloads.
Domain
AI Infrastructure / High-Performance Computing (HPC) / Data Center Networking
Deliverable
production ML models | infrastructure
Required skills
RDMA fabric design (InfiniBand/RoCE v2), BGP/EVPN-VXLAN routing, lossless Ethernet tuning, high-radix switching, network automation (Python/Ansible), cloud interconnects (Direct Connect/ExpressRoute), security segmentation, fabric validation and debugging.
Preferred skills
Dark fiber/DWDM operations, BlueField DPU/SmartNIC deployment, distributed training at 1000+ GPU scale, Kubernetes networking for GPU serving, expert certifications (CCIE/JNCIE/NVIDIA).
Technologies
InfiniBand (NDR/XDR), RoCE v2 (Spectrum-X/Tomahawk), ConnectX/BlueField NICs, 400/800G optics, Arista EOS, NVIDIA Cumulus, Cisco NX-OS, Junos, SONiC, Ansible, Nornir, NetBox, Prometheus, Grafana, AWS Direct Connect, Azure ExpressRoute.
Responsibilities
Architect greenfield network topologies for new GPU clusters; configure and validate lossless Ethernet and InfiniBand fabrics; manage routing protocols and network security; automate network operations and configuration; monitor fabric telemetry and debug performance issues; design cross-site and hybrid cloud connectivity.
Seniority
Senior, hands-on IC
