Senior Manager, Engineering - AI Inference
Core
Lead an engineering team to optimize large language model inference for speed, cost, and reliability in production environments.
Role type
Senior Engineering Manager (AI Inference)
Builds
Production-grade LLM serving stacks and optimized inference systems
Domain
AI Infrastructure / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
Engineering team leadership, LLM inference optimization, serving architecture design, performance profiling, CUDA/GPU internals, Python, C++, vLLM, SGLang
Preferred skills
CUDA development, Kubernetes, Docker, customer-facing technical solutions
Technologies
vLLM, SGLang, CUDA, Python, C++, Docker, Kubernetes
Responsibilities
Design and optimize serving architectures including prefill/decode disaggregation; Profile and tune deployments for latency, throughput, and cost; Partner with customers to move workloads from POC to production; Lead team delivery from experiments to production; Guide technical strategy and roadmaps with product and engineering leaders.
Seniority
Senior, hands-on IC with people management