Senior Network Site Reliability Engineer (NetSRE)
Core
Build and run the fundamental network infrastructure for a full-stack AI cloud platform, ensuring reliability, safety, and scalability for compute, storage, and networking services.
Role type
Senior Network Site Reliability Engineer (NetSRE)
Builds
High-availability network services, inter-site connectivity (DCI), and operational tooling for an AI cloud platform.
Domain
Cloud Infrastructure / AI / Networking
Deliverable
production ML models | infrastructure
Required skills
Linux system administration, network fundamentals (control/data plane, failure domains), high-availability system operations, software automation (Go/Python), infrastructure as code (IaC), CI/CD pipelines, observability (metrics/logs/traces), incident response and postmortems, change management workflows.
Preferred skills
High-throughput traffic processing (load balancers, tunneling, NAT64), low-level networking performance debugging (eBPF/XDP, DPDK, kernel internals), network-safe delivery pipelines, large-scale network telemetry analysis.
Technologies
Go, Python, Linux, eBPF, XDP, DPDK, IaC, CI/CD, container platforms.
Responsibilities
Define and own reliability goals (SLIs/SLOs, error budgets) for network services; drive reliability improvements across site readiness and inter-site connectivity; lead incident response and postmortems to implement durable fixes; build and evolve observability stacks for faster debugging; design safer change workflows including automation and canarying; collaborate with network engineers to embed operability into designs.
Seniority
Senior, hands-on IC