Technical Product Manager - AI Compute Platform
Core
Building the AI cloud platform that enables developers and enterprises to train state-of-the-art models and run production inference at scale, focusing on compute, storage, networking, and applied AI.
Role type
Senior Technical Product Manager (AI Compute Platform)
Builds
Full-stack AI cloud platform including GPU orchestration, cluster lifecycle management, reliability systems, and managed runtime platforms for frontier AI labs.
Domain
Cloud Infrastructure / AI Compute / HPC
Deliverable
production ML models | product features | infrastructure
Required skills
Product strategy and roadmap ownership, API contract design, cross-team execution across engineering and SRE, structured customer discovery and analytics, technical peer engagement with engineering leaders, success metric definition.
Preferred skills
Frontier AI customer experience, Kubernetes/Slurm/HPC environment familiarity, ML training/inference workflow expertise, NVIDIA stack and InfiniBand knowledge, cluster lifecycle management, observability product design, reliability engineering background.
Technologies
Kubernetes, Slurm, NVIDIA (CUDA, NCCL, DGX), InfiniBand, RoCE, Grafana, Datadog, vLLM, Ray, Token Factory
Responsibilities
Own end-to-end product responsibility for a specific platform slice including strategy, roadmap, and delivery; Design and own platform contracts (APIs, semantics, system events); Drive cross-team execution across platform engineering, networking, and support; Turn customer pain into product commitments through structured discovery; Engage engineering as a technical peer to debate design and trade-offs; Define and own success metrics based on customer and platform outcomes.
Seniority
Senior, hands-on IC