Network Engineer (Supercomputer Infrastructure)
Core
Design, build, and operate mission-critical, low-latency, high-bandwidth networks for AI supercomputer campuses, supporting GPU training and inference clusters.
Role type
Senior IC network engineer (supercomputer infrastructure)
Builds
AI training fabrics, inference front-ends, storage networks, and site/OT networks for GPU clusters
Domain
AI/HPC infrastructure, data center networking, supercomputer campuses
Deliverable
production ML models | infrastructure
Required skills
Layer 2/3 network design and troubleshooting, GitOps/IaC, network hardware procurement and deployment, network automation, root cause analysis, network documentation, cross-functional collaboration
Preferred skills
Cisco/Arista/Juniper/NVIDIA Spectrum-X switches, RoCEv2/InfiniBand, AI traffic patterns/NCCL, WDM/fiber plant management, QoS/multicast/redundancy protocols, network monitoring/telemetry, scripting (Bash/Python), Linux/Windows sysadmin
Technologies
Cisco, Arista, Juniper, NVIDIA Spectrum-X, RoCEv2, InfiniBand, GitOps, Terraform, Ansible, OTDR
Responsibilities
Design and implement highly available, low-latency networks for AI fabrics; maintain data center and campus networks; evaluate and deploy network hardware; contribute to network automation tooling; plan and coordinate network change windows; troubleshoot network issues affecting cluster health; provide direct networking support during cluster operations; create and update network documentation; collaborate on design issues and failure modes; perform job walks for new infrastructure; ensure compliance with cybersecurity standards
Seniority
Senior, hands-on IC
