Engineering Manager - Inference Performance
Core
Lead a team of inference performance engineers to optimize GPU workloads for high-performance AI inference, focusing on runtime efficiency, scheduling, and model deployment.
Role type
Senior Engineering Manager (Inference Performance)
Builds
High-performance inference runtime and engine components for AI models
Domain
AI Infrastructure / GPU Systems
Deliverable
production ML models
Required skills
GPU architecture and optimization, team leadership and hiring, technical roadmap execution, profiling and performance analysis, ML library familiarity (PyTorch, TensorRT), stakeholder alignment
Preferred skills
Inference engine experience (vLLM, SGLang), LLM optimization techniques, GPU kernel development (CUDA, Triton), startup scaling experience, hands-on systems engineering background
Technologies
PyTorch, TensorRT, TensorRT-LLM, CUDA, Triton, CUTLASS, vLLM, SGLang
Responsibilities
Lead, mentor, and grow a team of inference performance engineers; Hire top GPU and inference engineering talent; Own the technical roadmap for runtime performance work; Review designs and guide profiling/optimization efforts; Drive productionization of inference techniques (quantization, speculative decoding, KV-cache reuse); Turn performance wins into measurable outcomes (tokens/GPU-hour, latency, cost); Help bring up and tune new model architectures on new hardware; Partner with Infrastructure, Platform, and customer-facing teams to ship wins
Seniority
Senior, hands-on technical leadership (via careerplan.io/jobs/6b105131-d5d6-4d41-96ce-7fe71ff87769-engineering-manager-inference-performance-at-baseten)
