Site Reliability Engineer in Network Infrastructure
Core
Build and run the fundamental network infrastructure for a full-stack AI cloud platform, ensuring reliability and safety as the system scales.
Role type
Senior Site Reliability Engineer (Network Infrastructure)
Builds
High-availability network services, inter-site connectivity, and operational tooling for an AI cloud platform.
Domain
Cloud Infrastructure / Networking / AI Platform
Deliverable
production ML models | infrastructure
Required skills
Linux system administration, networking fundamentals (control/data plane, failure domains), incident response and postmortems, observability (metrics/logs/traces), automation and scripting (Go/Python), Infrastructure as Code (IaC), CI/CD pipelines.
Preferred skills
High-throughput traffic processing (load balancers, tunneling, NAT64), low-level networking performance debugging (eBPF, XDP, DPDK), network-safe delivery pipelines, large-scale network telemetry analysis.
Responsibilities
Define and own reliability goals (SLIs/SLOs, error budgets) for network services; drive reliability improvements across site readiness and inter-site connectivity; lead incident response and investigations; build and evolve observability stacks; design safer change workflows with automation and canarying; collaborate with network engineers to embed operability into designs.
Seniority
Senior, hands-on IC