CareerPlanSign in

Production Engineering Lead, Compute

San Francisco, CA💼 Full-time🗓 2026-07-21 → 2026-09-26

Core

Lead a production engineering team to operate a massive fleet of tens of thousands of GPUs, ensuring high availability and scaling compute infrastructure for AI customers.

Role type

Senior IC Production Engineering Lead (Compute Infrastructure)

Builds

Automated node lifecycle management, provisioning, health checks, and remediation tooling for GPU fleets.

Domain

AI Infrastructure / Data Center Operations / High-Performance Computing

Deliverable

production ML models | infrastructure

Required skills

SRE leadership, fleet availability management, automation engineering, on-call/escalation design, team hiring and growth

Preferred skills

GPU or HPC fleet experience, Kubernetes or Slurm, hardware failure analytics, customer-facing reliability

Technologies

Kubernetes, Slurm, GPU clusters

Responsibilities

Define SLOs and build tooling to move fleet availability numbers; build automation for node lifecycle without human touch; set on-call and escalation models; lead a team running large-scale GPU fleets.

Seniority

Senior, hands-on IC with leadership scope

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.