Production Engineering Lead, Compute
Core
Lead a production engineering team to operate a massive fleet of tens of thousands of GPUs, ensuring high availability and scaling compute infrastructure for AI customers.
Role type
Senior IC Production Engineering Lead (Compute Infrastructure)
Builds
Automated node lifecycle management, provisioning, health checks, and remediation tooling for GPU fleets.
Domain
AI Infrastructure / Data Center Operations / High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
SRE leadership, fleet availability management, automation engineering, on-call/escalation design, team hiring and growth
Preferred skills
GPU or HPC fleet experience, Kubernetes or Slurm, hardware failure analytics, customer-facing reliability
Technologies
Kubernetes, Slurm, GPU clusters
Responsibilities
Define SLOs and build tooling to move fleet availability numbers; build automation for node lifecycle without human touch; set on-call and escalation models; lead a team running large-scale GPU fleets.
Seniority
Senior, hands-on IC with leadership scope