CareerPlanSign in

Senior Product Manager – AI Inference Performance

US, CA, Santa Clara💼 Full-time💰 $208,000–$208,000🗓 2026-08-13 → 2026-09-25

Core

Own the inference performance roadmap for AI models and applications running on NVIDIA hardware, focusing on latency, efficiency, and cost per token across the entire inference stack.

Role type

Senior Product Manager (AI Inference Performance)

Builds

Platforms and capabilities for AI inference optimization, including agentic workloads, framework strategies, and benchmarking tooling.

Domain

AI/ML Infrastructure, GPU Computing, Inference Optimization

Deliverable

production ML models | product features

Required skills

AI inference optimization (KV caching, quantization, speculative decoding), inference framework strategy, product roadmap ownership, benchmarking methodology, release management, translating technical capabilities to business value

Preferred skills

Engineering experience with LLM inference profiling/optimization, open-source contributions to vLLM/SGLang/TensorRT-LLM, reading research papers for roadmap decisions

Technologies

TensorRT-LLM, vLLM, SGLang, NVIDIA Dynamo, Triton Inference Server

Responsibilities

Define performance strategy for agentic and multi-turn workloads; Partner with open-source communities and internal engineering teams; Define benchmark methodology and metrics (TTFT, ITL, throughput); Own release readiness, quality bars, and customer blocking issues

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.