Sr./Staff TPM - Inference Capacity
Core
Lead capacity planning and fleet strategy for the Inference Service organization to maximize utilization of one of the world's most advanced AI inference fleets.
Role type
Sr./Staff Technical Program Manager (Inference Capacity)
Builds
Rolling capacity models, datacenter bring-up plans, cluster allocation strategies, and internal capacity management tools.
Domain
AI Infrastructure / Large-scale ML Serving / Cloud Infrastructure
Deliverable
production ML models
Required skills
Capacity planning and forecasting, cross-functional program leadership, inference serving stack knowledge (model replicas, batching, KV cache), data fluency (SQL, Grafana, Python), incident management, stakeholder adoption of tools
Preferred skills
Experience with AI accelerator fleets (Habana, TPU, Inferentia, Trainium), experience with hyperscaler capacity planning
Technologies
Jira, Confluence, SQL, Grafana, Python, Flux
Responsibilities
Run weekly capacity planning and daily tracking with Engineering, product, and operations; drive capacity planning for new customer deployments and major model launches; lead planning around major infrastructure events; maintain Jira EPICs and Confluence pages for execution transparency; drive continuous improvement of capacity management platforms
Seniority
Senior, hands-on IC