Staff Network Production Engineer, Operations
Core
Own production reliability for global AI infrastructure network including edge, backbone, data center fabric, and GPU cluster interconnects.
Role type
Staff Network Production Operations Engineer
Builds
Global AI compute network infrastructure supporting thousands of GPUs
Domain
AI Infrastructure / Data Center Networking
Deliverable
production ML models | infrastructure
Required skills
Large-scale network operations, Incident response, Root cause analysis, Observability tooling, Operational automation, SLI/SLO definition, Technical mentorship
Preferred skills
NVIDIA/Mellanox platforms, Traffic analysis tools (Kentik/Arbor), Post-incident learning programs
Technologies
Python, Streaming telemetry, SNMP, NetFlow/sFlow, Grafana, Prometheus, ThousandEyes, RDMA/RoCE, PFC, ECN, DCQCN, BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, TCP/IP, Arista EOS, Juniper Junos
Responsibilities
Own uptime across global edge, backbone, data center, and GPU cluster network; Lead end-to-end response for high-severity network events; Drive RCAs for production incidents and author remediation plans; Improve network monitoring stack using streaming telemetry and tools like Kentik and Grafana; Author and maintain runbooks, escalation playbooks, and SOPs; Write Python-based tooling to automate remediation workflows; Partner with Architecture and SRE teams to define and track network reliability metrics.
Seniority
Staff, hands-on IC with mentorship responsibilities