CareerPlanGet AI match score →

Reliability Engineer, Supercomputing

San Francisco💼 Full-time💰 $350,000–$350,000🗓 2026-06-24 → 2026-07-31

Core

Ensuring the reliability of GPU supercomputing fleets by diagnosing hardware issues, managing drivers/firmware, and collaborating with vendors to maintain system stability for AI research.

Role type

Reliability Engineer (Supercomputing)

Builds

Reliable GPU clusters for frontier AI research

Domain

AI / Supercomputing / Hardware Reliability

Deliverable

infrastructure

Required skills

Python or Rust, Linux systems, Kubernetes or Slurm, hardware debugging, vendor engagement, firmware lifecycle management, postmortem writing

Preferred skills

Linux kernel literacy, statistical rigor in reliability analysis, GPU hardware health expertise (Xid errors, NVLink, DCGM), out-of-band management (BMC, iDRAC, IPMI, Redfish), hardware reliability ownership at scale

Technologies

Python, Rust, Kubernetes, Slurm, Linux, BMC, iDRAC, IPMI, Redfish, DCGM, NVLink

Responsibilities

Investigate and remediate issues across large GPU clusters; Own drivers, kernel surface, and diagnostics spanning hardware, firmware, and OS; Automate fleet reliability monitoring and error rate analysis; Drive firmware lifecycle including qualification and regression analysis; Engage hardware vendors directly to resolve issues and manage RMA flows; Monitor GPU hardware health signals and implement reliability improvements; Write postmortems and vendor cases

Seniority

Individual Contributor, hands-on engineering

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗