CareerPlanSign in

Staff Network Production Engineer, Operations

San Francisco, CA - US💼 Full-time💰 $195,000–$195,000🗓 2026-08-13 → 2026-09-26

Core

Own production reliability for global AI infrastructure network including edge, backbone, data center fabric, and GPU cluster interconnects.

Role type

Staff Network Production Operations Engineer

Builds

Global AI compute network infrastructure supporting thousands of GPUs

Domain

AI Infrastructure / Data Center Networking

Deliverable

production ML models | infrastructure

Required skills

Large-scale network operations, Incident response, Root cause analysis, Observability tooling, Operational automation, SLI/SLO definition, Technical mentorship

Preferred skills

NVIDIA/Mellanox platforms, Traffic analysis tools (Kentik/Arbor), Post-incident learning programs

Technologies

Python, Streaming telemetry, SNMP, NetFlow/sFlow, Grafana, Prometheus, ThousandEyes, RDMA/RoCE, PFC, ECN, DCQCN, BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, TCP/IP, Arista EOS, Juniper Junos

Responsibilities

Own uptime across global edge, backbone, data center, and GPU cluster network; Lead end-to-end response for high-severity network events; Drive RCAs for production incidents and author remediation plans; Improve network monitoring stack using streaming telemetry and tools like Kentik and Grafana; Author and maintain runbooks, escalation playbooks, and SOPs; Write Python-based tooling to automate remediation workflows; Partner with Architecture and SRE teams to define and track network reliability metrics.

Seniority

Staff, hands-on IC with mentorship responsibilities

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.