CareerPlanSign in

Software Engineer, Hardware Health

San Francisco💼 Full-time🗓 2026-05-11 → 2026-09-26

Core

Build critical infrastructure to observe, detect, remediate, and verify hardware issues across millions of GPUs, CPUs, and networking components to ensure reliable compute for large-scale AI training and inference.

Role type

Senior IC infrastructure engineer (hardware health & observability)

Builds

Automated remediation systems, health checks, and global cluster management tooling

Domain

Hyperscale AI compute infrastructure, distributed systems, hardware lifecycle management

Deliverable

production ML models | infrastructure

Required skills

Python, shell scripting, large-scale distributed systems, SQL, PromQL, systems debugging, operational tooling

Preferred skills

low-level hardware systems (PCIe, InfiniBand, RoCE), Linux kernel tuning, GPU cluster operations, automated remediation, fleet lifecycle management

Technologies

Python, shell, SQL, PromQL, Linux, PCIe, InfiniBand, RoCE

Responsibilities

Define and maintain health signals across GPUs, CPUs, and networking; build and evolve health checks for failure detection and remediation; investigate hardware failures and system-level issues; own node lifecycle workflows (drain, quarantine, repair, RMA); build automation for global cluster management; partner with reliability and provider teams to integrate health signals.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.