CareerPlanSign in

Senior Manager, Engineering - AI Inference

San Francisco, CA - US💼 Full-time💰 $250,000–$250,000🗓 2026-09-23 → 2026-09-25

Core

Lead an engineering team to optimize large language model inference for speed, cost, and reliability in production environments.

Role type

Senior Engineering Manager (AI Inference)

Builds

Production-grade LLM serving stacks and optimized inference systems

Domain

AI Infrastructure / Large Language Models / High-Performance Computing

Deliverable

production ML models

Required skills

Engineering team leadership, LLM inference optimization, serving architecture design, performance profiling, CUDA/GPU internals, Python, C++, vLLM, SGLang

Preferred skills

CUDA development, Kubernetes, Docker, customer-facing technical solutions

Technologies

vLLM, SGLang, CUDA, Python, C++, Docker, Kubernetes

Responsibilities

Design and optimize serving architectures including prefill/decode disaggregation; Profile and tune deployments for latency, throughput, and cost; Partner with customers to move workloads from POC to production; Lead team delivery from experiments to production; Guide technical strategy and roadmaps with product and engineering leaders.

Seniority

Senior, hands-on IC with people management

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.