CareerPlanSign in

GPU Cluster Engineer, Networking

San Francisco, CA💼 Full-time🗓 2026-08-06 → 2026-09-25

Core

Design, build, and operate the complete networking stack for large-scale GPU clusters powering multimodal AI models, including RDMA fabrics, data center perimeters, and cross-site connectivity.

Role type

Senior Network Engineer (AI Infrastructure)

Builds

High-performance GPU cluster fabrics (InfiniBand/RoCE), data center networks, and hybrid cloud interconnects for frontier AI workloads.

Domain

AI Infrastructure / High-Performance Computing (HPC) / Data Center Networking

Deliverable

production ML models | infrastructure

Required skills

RDMA fabric design (InfiniBand/RoCE v2), BGP/EVPN-VXLAN routing, lossless Ethernet tuning, high-radix switching, network automation (Python/Ansible), cloud interconnects (Direct Connect/ExpressRoute), security segmentation, fabric validation and debugging.

Preferred skills

Dark fiber/DWDM operations, BlueField DPU/SmartNIC deployment, distributed training at 1000+ GPU scale, Kubernetes networking for GPU serving, expert certifications (CCIE/JNCIE/NVIDIA).

Technologies

InfiniBand (NDR/XDR), RoCE v2 (Spectrum-X/Tomahawk), ConnectX/BlueField NICs, 400/800G optics, Arista EOS, NVIDIA Cumulus, Cisco NX-OS, Junos, SONiC, Ansible, Nornir, NetBox, Prometheus, Grafana, AWS Direct Connect, Azure ExpressRoute.

Responsibilities

Architect greenfield network topologies for new GPU clusters; configure and validate lossless Ethernet and InfiniBand fabrics; manage routing protocols and network security; automate network operations and configuration; monitor fabric telemetry and debug performance issues; design cross-site and hybrid cloud connectivity.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.