Senior Technical Program Manager, Cluster Operations & Quota Management
Core
Own the end-to-end operating system for quota and cluster execution, turning allocation decisions into usable GPU capacity across a fleet.
Role type
Senior Technical Program Manager (Infrastructure/Compute)
Builds
Quota implementation machinery, cluster move playbooks, and cross-team coordination processes for GPU fleet capacity.
Domain
High-performance computing (HPC) / GPU infrastructure / Cloud platform operations
Deliverable
production ML models | infrastructure
Required skills
End-to-end program ownership, cross-functional coordination without authority, dependency management, process improvement, automation advocacy, stakeholder convening, change management, capacity planning, vendor coordination, validation of environment readiness
Preferred skills
Experience in compute-intensive environments, background in infrastructure or platform engineering, ability to drive operating model changes
Technologies
GPU fleet management tools, cluster orchestration systems, quota management platforms
Responsibilities
Prepare and sequence quota and cluster changes in advance; run cluster moves and cycle cutovers cleanly; coordinate provisioning with infrastructure, platform, HPC, and vendor teams; validate identity, access, and utilization for researchers; chase dependencies and surface blockers early; streamline handoffs between teams; convene working groups to fix root causes in the operating model.
Seniority
Senior, hands-on IC