CareerPlanSign in

Production Engineer, Compute

San Francisco, CA💼 Full-time🗓 2026-06-16 → 2026-09-26

Core

Building automation, observability, and repair pipelines for a hyperscale GPU fleet to ensure reliability and throughput for frontier AI compute.

Role type

Senior IC production engineer (compute infrastructure)

Builds

Automated repair pipelines, GPU qualification platforms, fleet observability layers, and Redfish/BMC tooling.

Domain

AI infrastructure, data center operations, hardware lifecycle management

Deliverable

production ML models | infrastructure

Required skills

Hardware failure mode analysis, firmware/silicon level reasoning, automation pipeline design, incident management, AI tooling fluency (LLM APIs, agentic frameworks), Go or Python

Preferred skills

BMC/Redfish/IPMI tooling, GPU burn-in frameworks, workflow orchestration (Temporal, Cadence), metrics pipelines (Prometheus, Grafana)

Technologies

Kubernetes, Redfish, BMC, Prometheus, Grafana, Temporal, Cadence, LLM APIs

Responsibilities

Own compute fleet health end-to-end including metrics pipelines and alerting; Turn deployment/repair into automated pipelines; Design and expand GPU qualification platforms; Own Redfish and BMC tooling for firmware-level telemetry; Ensure end-to-end reliability and scalability of the compute fleet at scale

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.