Product Manager, Compute Platform
Core
Build scheduling, orchestration, and capacity management systems for Anthropic's compute infrastructure to power model training, evaluation, and inference workloads.
Role type
Senior Product Manager, Compute Platform
Builds
Job scheduling primitives, capacity allocation policies, observability tooling, and resource management systems for GPU/accelerator clusters.
Domain
AI/ML Infrastructure, High-Performance Computing (HPC), Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Product management for technical infrastructure, distributed systems knowledge, job scheduling/orchestration expertise, stakeholder management across Research/Finance/Engineering, trade-off analysis (utilization vs cost vs latency), capacity planning, observability strategy.
Preferred skills
Experience with Kubernetes/Slurm/Borg/YARN, GPU/accelerator scheduling (gang-scheduling, topology-aware placement), SLA definition for compute workloads, cloud/on-prem capacity modeling, hypergrowth environment experience.
Technologies
Kubernetes, Slurm, Borg, YARN, GPU clusters
Responsibilities
Define semantic layer for job scheduling (abstractions, priority tiers, preemption policies), partner with engineering to design scheduling capabilities maximizing cluster utilization, drive roadmap for capacity management and fairness policies, build observability dashboards for cluster health and resource waste, collaborate on capacity planning and cost-to-serve analytics.
Seniority
Senior, hands-on IC