Product Manager, Inference Platform
Core
Define the product strategy and roadmap for a mission-critical inference platform that enables AI companies to deploy, scale, and manage large language models reliably across multi-cloud environments.
Role type
Senior Product Manager, Inference Platform
Builds
A unified platform for production AI inference, including autoscaling, traffic routing, failover, and release management.
Domain
Cloud Infrastructure / AI Model Serving
Deliverable
production ML models
Required skills
Product management for infrastructure/distributed systems, end-to-end ownership (backend to UX), cross-team roadmap driving, defining new categories, scaling/routing/failover reasoning
Preferred skills
GPU infrastructure, Kubernetes, serving frameworks (vLLM, TensorRT-LLM, SGLang)
Technologies
Kubernetes, vLLM, TensorRT-LLM, SGLang, Multi-cloud
Responsibilities
Own workload scaling policies (autoscaling, placement, compliance), ensure production reliability (traffic routing, failover, health recovery), build release engines (canary, shadow, A/B), drive cost/performance optimization, reduce MTTR via self-serve incident management, set the roadmap for infrastructure teams.
Seniority
Senior, hands-on IC