CareerPlanSign in

Engineering Manager - Inference Performance

San Francisco💼 Full-time🗓 2026-10-07

Core

Lead a team of inference performance engineers to optimize GPU workloads for high-performance AI inference, focusing on runtime efficiency, scheduling, and model deployment.

Role type

Senior Engineering Manager (Inference Performance)

Builds

High-performance inference runtime and engine components for AI models

Domain

AI Infrastructure / GPU Systems

Deliverable

production ML models

Required skills

GPU architecture and optimization, team leadership and hiring, technical roadmap execution, profiling and performance analysis, ML library familiarity (PyTorch, TensorRT), stakeholder alignment

Preferred skills

Inference engine experience (vLLM, SGLang), LLM optimization techniques, GPU kernel development (CUDA, Triton), startup scaling experience, hands-on systems engineering background

Technologies

PyTorch, TensorRT, TensorRT-LLM, CUDA, Triton, CUTLASS, vLLM, SGLang

Responsibilities

Lead, mentor, and grow a team of inference performance engineers; Hire top GPU and inference engineering talent; Own the technical roadmap for runtime performance work; Review designs and guide profiling/optimization efforts; Drive productionization of inference techniques (quantization, speculative decoding, KV-cache reuse); Turn performance wins into measurable outcomes (tokens/GPU-hour, latency, cost); Help bring up and tune new model architectures on new hardware; Partner with Infrastructure, Platform, and customer-facing teams to ship wins