Site Reliability Engineer, Compute
Core
Build automation, observability, and repair pipelines for a hyperscale GPU fleet to ensure compute availability and reliability for frontier AI models.
Role type
Senior IC Site Reliability Engineer (Compute Infrastructure)
Builds
Automated repair workflows, GPU qualification platforms, fleet observability layers, and Redfish/BMC tooling for 10s to 100s of GWs of compute.
Domain
AI Infrastructure / Hardware-Software Convergence / Data Center Operations
Deliverable
production ML models | infrastructure
Required skills
GPU failure mode analysis, firmware/silicon-level debugging, Kubernetes orchestration, Redfish/BMC tooling, automation pipeline design, incident management, AI tooling fluency (LLM APIs, agentic frameworks)
Preferred skills
Hardware lifecycle management, RMA automation, burn-in frameworks, workflow engines (Temporal, Cadence), metrics pipelines (Prometheus, Grafana), Go or Python
Technologies
Kubernetes, Redfish, BMC, IPMI, Prometheus, Grafana, Temporal, Cadence, LLM APIs, MCP servers
Responsibilities
Own compute fleet health end-to-end including metrics pipelines and alerting; Turn deployment/repair into automated pipelines; Design and expand GPU qualification platforms; Own Redfish and BMC tooling for firmware telemetry; Ensure end-to-end reliability and scalability of the compute fleet at scale.
Seniority
Senior, hands-on IC