Senior Product Manager – AI Inference Performance
Core
Own the inference performance roadmap for AI models and applications running on NVIDIA hardware, focusing on latency, efficiency, and cost per token across the entire inference stack.
Role type
Senior Product Manager (AI Inference Performance)
Builds
Platforms and capabilities for AI inference optimization, including agentic workloads, framework strategies, and benchmarking tooling.
Domain
AI/ML Infrastructure, GPU Computing, Inference Optimization
Deliverable
production ML models | product features
Required skills
AI inference optimization (KV caching, quantization, speculative decoding), inference framework strategy, product roadmap ownership, benchmarking methodology, release management, translating technical capabilities to business value
Preferred skills
Engineering experience with LLM inference profiling/optimization, open-source contributions to vLLM/SGLang/TensorRT-LLM, reading research papers for roadmap decisions
Technologies
TensorRT-LLM, vLLM, SGLang, NVIDIA Dynamo, Triton Inference Server
Responsibilities
Define performance strategy for agentic and multi-turn workloads; Partner with open-source communities and internal engineering teams; Define benchmark methodology and metrics (TTFT, ITL, throughput); Own release readiness, quality bars, and customer blocking issues
Seniority
Senior, hands-on IC