Network Operations Engineer, AI Networking
Core
Operate and improve large-scale Ethernet fabrics supporting GPU clusters, storage systems, and management infrastructure for AI training and inference workloads.
Role type
Senior IC network operations engineer (AI networking)
Builds
Highly available GPU infrastructure for AI training and inference
Domain
AI infrastructure / Data center networking
Deliverable
production ML models | infrastructure
Required skills
Layer 2/3 networking, BGP, OSPF, ECMP, MLAG, LACP, VRFs, VLANs, physical-layer troubleshooting (fiber optics, transceivers), hardware lifecycle management, root-cause analysis, network automation (Python), observability (telemetry, dashboards)
Preferred skills
AI/HPC network environments, NVIDIA AI networking technologies, RoCE v2, RDMA, lossless Ethernet, 100G/200G/400G/800G Ethernet, GPU platforms (NVIDIA HGX, DGX, GB200), distributed storage (VAST, DDN), cloud provider operations (AWS, Azure, GCP)
Technologies
Cisco NX-OS, Arista EOS, NVIDIA Spectrum / Cumulus Linux, Juniper JunOS, Prometheus, Grafana, gNMI, Terraform, Git, REST APIs
Responsibilities
Monitor and resolve network incidents to meet SLOs, execute production network changes and maintenance, manage hardware lifecycle and replacements, support new AI cluster deployments and migrations, build monitoring and alerting systems, develop operational runbooks and automate repetitive tasks
Seniority
Senior, hands-on IC