Reliability Engineer, Supercomputing
Core
Ensuring the reliability of GPU supercomputing fleets by diagnosing hardware issues, managing drivers/firmware, and collaborating with vendors to maintain system stability for AI research.
Role type
Reliability Engineer (Supercomputing)
Builds
Reliable GPU clusters for frontier AI research
Domain
AI / Supercomputing / Hardware Reliability
Deliverable
infrastructure
Required skills
Python or Rust, Linux systems, Kubernetes or Slurm, hardware debugging, vendor engagement, firmware lifecycle management, postmortem writing
Preferred skills
Linux kernel literacy, statistical rigor in reliability analysis, GPU hardware health expertise (Xid errors, NVLink, DCGM), out-of-band management (BMC, iDRAC, IPMI, Redfish), hardware reliability ownership at scale
Technologies
Python, Rust, Kubernetes, Slurm, Linux, BMC, iDRAC, IPMI, Redfish, DCGM, NVLink
Responsibilities
Investigate and remediate issues across large GPU clusters; Own drivers, kernel surface, and diagnostics spanning hardware, firmware, and OS; Automate fleet reliability monitoring and error rate analysis; Drive firmware lifecycle including qualification and regression analysis; Engage hardware vendors directly to resolve issues and manage RMA flows; Monitor GPU hardware health signals and implement reliability improvements; Write postmortems and vendor cases
Seniority
Individual Contributor, hands-on engineering