Global Capacity Manager
Core
Architecting, securing, and optimizing the global GPU fleet to power AI workloads for customers, ensuring 99.9% uptime and elite unit economics.
Role type
Senior IC infrastructure engineer (GPU fleet orchestration)
Builds
Multi-cloud capacity management systems, automated GPU operators, and infrastructure readiness for next-gen hardware (NVIDIA Blackwell)
Domain
Cloud infrastructure + AI hardware supply chain
Deliverable
production ML models
Required skills
Kubernetes (taints, cordons, node draining), Go, Python, financial modeling, GPU lifecycle management, multi-cloud orchestration
Preferred skills
Experience with hyperscalers (GCP, AWS, Azure), specialized GPU provider experience, NVIDIA Blackwell architecture knowledge
Technologies
Kubernetes, Go, Python, NVIDIA Blackwell (B200), H100
Responsibilities
Lead specialized GPU pods managing acquisition and maintenance; execute complex workload migrations and deployment drains; design scalable capacity management systems; model ROI for GPU spend; partner with SRE/Infra teams; lead incident response for capacity outages
Seniority
Senior, hands-on IC