Software Engineer, Hardware Health
Core
Build critical infrastructure to observe, detect, remediate, and verify hardware issues across millions of GPUs, CPUs, and networking components to ensure reliable compute for large-scale AI training and inference.
Role type
Senior IC infrastructure engineer (hardware health & observability)
Builds
Automated remediation systems, health checks, and global cluster management tooling
Domain
Hyperscale AI compute infrastructure, distributed systems, hardware lifecycle management
Deliverable
production ML models | infrastructure
Required skills
Python, shell scripting, large-scale distributed systems, SQL, PromQL, systems debugging, operational tooling
Preferred skills
low-level hardware systems (PCIe, InfiniBand, RoCE), Linux kernel tuning, GPU cluster operations, automated remediation, fleet lifecycle management
Technologies
Python, shell, SQL, PromQL, Linux, PCIe, InfiniBand, RoCE
Responsibilities
Define and maintain health signals across GPUs, CPUs, and networking; build and evolve health checks for failure detection and remediation; investigate hardware failures and system-level issues; own node lifecycle workflows (drain, quarantine, repair, RMA); build automation for global cluster management; partner with reliability and provider teams to integrate health signals.
Seniority
Senior, hands-on IC