Production Engineer, Compute
Core
Building automation, observability, and repair pipelines for a hyperscale GPU fleet to ensure reliability and throughput for frontier AI compute.
Role type
Senior IC production engineer (compute infrastructure)
Builds
Automated repair pipelines, GPU qualification platforms, fleet observability layers, and Redfish/BMC tooling.
Domain
AI infrastructure, data center operations, hardware lifecycle management
Deliverable
production ML models | infrastructure
Required skills
Hardware failure mode analysis, firmware/silicon level reasoning, automation pipeline design, incident management, AI tooling fluency (LLM APIs, agentic frameworks), Go or Python
Preferred skills
BMC/Redfish/IPMI tooling, GPU burn-in frameworks, workflow orchestration (Temporal, Cadence), metrics pipelines (Prometheus, Grafana)
Technologies
Kubernetes, Redfish, BMC, Prometheus, Grafana, Temporal, Cadence, LLM APIs
Responsibilities
Own compute fleet health end-to-end including metrics pipelines and alerting; Turn deployment/repair into automated pipelines; Design and expand GPU qualification platforms; Own Redfish and BMC tooling for firmware-level telemetry; Ensure end-to-end reliability and scalability of the compute fleet at scale
Seniority
Senior, hands-on IC