CareerPlanSign in

Senior Network Production Operations Engineer

San Francisco, CA - US💼 Full-time💰 $165,000–$165,000🗓 2026-09-29 → 2026-09-30

Core

Support production reliability for global AI infrastructure including edge, backbone, data center fabric, and GPU cluster interconnects.

Role type

Senior Network Production Operations Engineer

Builds

Global AI network infrastructure (edge, backbone, data center, GPU clusters)

Domain

AI Infrastructure / Data Center Networking

Deliverable

production ML models

Required skills

Python scripting, BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, TCP/IP, RDMA/RoCE, PFC, ECN, DCQCN, Arista EOS, Juniper Junos, multi-vendor leaf-spine CLOS, observability tooling (Grafana, Prometheus, ThousandEyes, Kentik), incident response, root cause analysis, automation of remediation workflows

Preferred skills

NVIDIA/Mellanox networking platforms, Kentik or Arbor traffic analysis, SLI/SLO definition, operating 10K+ device fleets

Technologies

Python, Arista EOS, Juniper Junos, Grafana, Prometheus, ThousandEyes, Kentik, SNMP, NetFlow, sFlow, RDMA, RoCE

Responsibilities

Support uptime across global edge, backbone, data center, and GPU cluster networks; write and maintain Python-based automation to reduce operational toil; participate in high-severity network events for detection, triage, and mitigation; perform root cause analysis using monitoring tooling; build automation on top of streaming telemetry and monitoring stacks; maintain runbooks, escalation playbooks, and SOPs; help build automated dashboards and alerting for reliability metrics

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 898,000+ jobs from 20+ sources.