CareerPlanGet AI match score →

Site Reliability Engineer - Ops & Automation

US and Canada Offices💼 Full-time🗓 2025-10-14 → 2026-07-31

Core

Building a high-performance SRE function to support one of the world's fastest-growing AI inference services powered by the Wafer-Scale Engine (WSE), delivering infrastructure for frontier-class models.

Role type

Senior IC Site Reliability Engineer (Ops & Automation)

Builds

Robust continuous delivery pipelines, self-service capabilities, and high-stakes production environments for AI inference.

Domain

AI Infrastructure / High-performance Computing

Deliverable

production ML models

Required skills

SRE operations, Kubernetes, Python or Go, Prometheus, Grafana, observability, capacity planning, reliability practices (SLOs, post-mortems)

Preferred skills

GitOps (Argo CD / Flux), continuous delivery pipeline development, Bazel, on-prem or multi-datacenter environments

Technologies

Kubernetes, Bazel, Prometheus, Grafana, InfluxDB, Python, Go, Argo CD, Flux

Responsibilities

Execute operational tasks including releases, capacity changes, and cluster upgrades; develop self-service CD pipelines; build reusable automation and internal developer tools; develop telemetry, observability, and alerting solutions; collaborate on high-impact automation opportunities; contribute to reliability practices.

Seniority

Mid-level (2-4+ years), hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗