Technical Program Manager, Compute
Core
Drive planning, coordination, and execution of programs to keep Anthropic's compute infrastructure running efficiently at scale, managing supply, capacity allocation, and utilization across the compute lifecycle.
Role type
Senior Technical Program Manager (Compute Infrastructure)
Builds
Operational visibility, coordination mechanisms, and processes for the compute fleet supporting model training, evaluation, and inference.
Domain
AI/ML Infrastructure, Cloud Computing, High-Performance Computing
Deliverable
infrastructure
Required skills
Technical program management, cross-functional coordination, cloud infrastructure knowledge, resource orchestration, stakeholder management, process design, trade-off analysis, operational gap identification
Preferred skills
Multi-cloud management (AWS, GCP, Azure), job scheduling systems (Kubernetes, Slurm, Borg, YARN), GPU/accelerator infrastructure, observability tooling, capacity planning, hypergrowth scaling experience
Technologies
Kubernetes, Slurm, Borg, YARN, AWS, GCP, Azure, GPU clusters
Responsibilities
Own and drive critical programs across the compute lifecycle; Build and maintain operational visibility into the compute fleet; Lead cross-functional coordination for compute transitions; Partner with engineering and research leadership to align on resource planning; Identify and close operational gaps across the compute pipeline; Own trade-off discussions between utilization, cost, latency, and reliability; Develop and improve processes and frameworks for planning and executing compute programs
Seniority
Senior, hands-on IC