Staff Site Reliability Engineer – Automation and Platform
Core
Lead engineering efforts to eliminate toil at scale by driving self-service delivery pipelines, shared observability tooling, and GitOps-driven CD for model releases and cluster upgrades for ultra-reliable AI inference infrastructure.
Role type
Staff Site Reliability Engineer (Platform & Automation)
Builds
Self-service platforms, internal tooling, declarative GitOps-driven CI/CD systems, and shared observability common tooling for AI inference workloads.
Domain
AI/ML inference systems, Wafer-Scale Engine (WSE) clusters, multi-datacenter and cloud-based solutions.
Deliverable
production ML models | infrastructure
Required skills
SRE/infrastructure engineering at large scale, operating large scale heterogeneous clusters with proprietary control planes, designing and delivering CI/CD or GitOps systems (Argo CD), hands-on observability (Loki, Tempo, Mimir, Prometheus), leading complex end-to-end projects, influencing cross-functional stakeholders.
Preferred skills
Bazel or large-scale build systems, AI/ML inference systems and model serving runtimes, predictive autoscaling, chaos engineering, cost-aware capacity planning.
Technologies
Argo CD, Loki, Tempo, Mimir, Prometheus, Bazel, GitOps, WSE.
Responsibilities
Define and implement robust strategies for delivering and running software reliably across multiple datacenters and cloud solutions; Architect self-service platforms and internal tooling for critical workflows; Define and evolve reliability practices including SLOs, SLIs, error budgets, blameless postmortems, and chaos testing; Mentor mid-level SREs and support critical incident escalations; Measure and drive impact through metrics like toil reduction, deployment velocity, and SLO compliance.
Seniority
Staff, hands-on IC with mentorship and strategy