Director, Product Management – AI Inference Platform
Core
Lead strategy and execution for a next-generation AI Inference Platform, defining how modern AI models are executed, served, and optimized at scale across distributed compute environments.
Role type
Director / Principal Product Manager (AI Infrastructure)
Builds
Large-scale AI inference systems, compute execution, inference serving, and control plane services
Domain
AI/ML Infrastructure, Cloud Platforms, Distributed Systems
Deliverable
production ML models
Required skills
Product management in AI/ML infrastructure, Large-scale serving architectures, LLM inference concepts (prefill/decode/KV cache), Scalable high-performance platform delivery, Hardware-software co-design
Preferred skills
Accelerator-based systems optimization, Modern inference frameworks (TensorRT-LLM, vLLM), Production-scale AI systems
Technologies
TensorRT-LLM, vLLM
Responsibilities
Define product vision and roadmap for AI inference systems, Drive inference architecture and orchestration strategy, Partner with engineering for hardware-software co-design, Shape platform capabilities for latency/throughput/cost efficiency, Own and evolve control plane services, Define and track key performance metrics (P99 latency, TTFT, throughput)
Seniority
Director, hands-on IC with strategic leadership