Software Engineer, Compute (GPU)
Core
Build automation pipelines for GPU fleet health, repair, and qualification to support 10-100s of GWs of compute infrastructure.
Role type
Senior IC software engineer (GPU infrastructure & automation)
Builds
Automated repair pipelines, GPU qualification platforms, fleet observability, and Redfish/BMC tooling
Domain
AI infrastructure, data center operations, hardware lifecycle management
Deliverable
production ML models | infrastructure
Required skills
Go, Python, Kubernetes, Redfish, BMC, firmware-level debugging, incident management, automation engineering
Preferred skills
Hardware lifecycle management, RMA automation, GPU qualification frameworks, workflow orchestration (Temporal, Cadence), metrics pipelines (Prometheus, Grafana)
Technologies
Kubernetes, Redfish, BMC, Prometheus, Grafana, Temporal, Cadence
Responsibilities
Own compute fleet health end-to-end including metrics pipelines and alerting; Turn deployment/repair into automated pipelines; Design and expand GPU qualification platforms; Own Redfish and BMC tooling for firmware-level telemetry; Ensure end-to-end reliability and scalability of the compute fleet at scale
Seniority
Senior, hands-on IC