CareerPlanSign in

Site Reliability Engineer, Compute

San Francisco, CA💼 Full-time🗓 2026-07-20 → 2026-09-26

Core

Build automation, observability, and repair pipelines for a hyperscale GPU fleet to ensure compute availability and reliability for frontier AI models.

Role type

Senior IC Site Reliability Engineer (Compute Infrastructure)

Builds

Automated repair workflows, GPU qualification platforms, fleet observability layers, and Redfish/BMC tooling for 10s to 100s of GWs of compute.

Domain

AI Infrastructure / Hardware-Software Convergence / Data Center Operations

Deliverable

production ML models | infrastructure

Required skills

GPU failure mode analysis, firmware/silicon-level debugging, Kubernetes orchestration, Redfish/BMC tooling, automation pipeline design, incident management, AI tooling fluency (LLM APIs, agentic frameworks)

Preferred skills

Hardware lifecycle management, RMA automation, burn-in frameworks, workflow engines (Temporal, Cadence), metrics pipelines (Prometheus, Grafana), Go or Python

Technologies

Kubernetes, Redfish, BMC, IPMI, Prometheus, Grafana, Temporal, Cadence, LLM APIs, MCP servers

Responsibilities

Own compute fleet health end-to-end including metrics pipelines and alerting; Turn deployment/repair into automated pipelines; Design and expand GPU qualification platforms; Own Redfish and BMC tooling for firmware telemetry; Ensure end-to-end reliability and scalability of the compute fleet at scale.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.