Senior Network Production Operations Engineer
Core
Support production reliability for global AI infrastructure including edge, backbone, data center fabric, and GPU cluster interconnects.
Role type
Senior Network Production Operations Engineer
Builds
Global AI network infrastructure (edge, backbone, data center, GPU clusters)
Domain
AI Infrastructure / Data Center Networking
Deliverable
production ML models
Required skills
Python scripting, BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, TCP/IP, RDMA/RoCE, PFC, ECN, DCQCN, Arista EOS, Juniper Junos, multi-vendor leaf-spine CLOS, observability tooling (Grafana, Prometheus, ThousandEyes, Kentik), incident response, root cause analysis, automation of remediation workflows
Preferred skills
NVIDIA/Mellanox networking platforms, Kentik or Arbor traffic analysis, SLI/SLO definition, operating 10K+ device fleets
Technologies
Python, Arista EOS, Juniper Junos, Grafana, Prometheus, ThousandEyes, Kentik, SNMP, NetFlow, sFlow, RDMA, RoCE
Responsibilities
Support uptime across global edge, backbone, data center, and GPU cluster networks; write and maintain Python-based automation to reduce operational toil; participate in high-severity network events for detection, triage, and mitigation; perform root cause analysis using monitoring tooling; build automation on top of streaming telemetry and monitoring stacks; maintain runbooks, escalation playbooks, and SOPs; help build automated dashboards and alerting for reliability metrics
Seniority
Senior, hands-on IC