Capacity Operations Manager
Core
Own the operational and analytical supply side of the GPU fleet to maximize healthy compute capacity across neocloud and bare metal environments.
Role type
Senior IC Capacity Operations Manager (GPU Fleet)
Builds
GPU fleet lifecycle, health, observability, and utilization monitoring
Domain
AI Infrastructure / GPU Compute / Cloud Operations
Deliverable
production ML models
Required skills
GPU fleet lifecycle management, supplier relationship management, capacity reconciliation, SLA monitoring, data analysis, cross-functional collaboration
Preferred skills
Hyperscaler or neo cloud provider experience, NVIDIA H100/H200/GB200 hardware lifecycle knowledge, supply chain dynamics expertise, formal supplier corrective actions experience
Technologies
Neocloud, Bare metal, NVIDIA H100/H200/GB200
Responsibilities
Drive suppliers to maintain maximum GPU fleet online and healthy status; Maintain live reconciliation of contracted vs. provisioned vs. healthy vs. utilized capacity; Own replacement SLAs, MTTR, and RMA cycle times for suppliers; Monitor SLA performance, file credit claims, and drive remediation plans; Coordinate internal communications regarding supplier maintenance impacting availability
Seniority
Senior, hands-on IC