CareerPlanSign in

Engineering Manager, Deep Learning Inference

6 Locations💼 Full-time💰 $184,000–$184,000🗓 2026-08-05 → 2026-09-26

Core

Lead a world-class engineering team developing and optimizing open-source deep learning inference frameworks (vLLM, SGLang, FlashInfer) for NVIDIA GPUs to enable scalable, real-time AI model deployment.

Role type

Senior Engineering Manager, Deep Learning Inference Software

Builds

Open-source inference frameworks and optimized inference pipelines for LLMs and multimodal generative AI

Domain

AI/ML Infrastructure, GPU Computing, High-Performance Computing

Deliverable

production ML models

Required skills

Technical leadership, C/C++ software design, GPU programming (CUDA, Triton, CUTLASS), performance optimization, multi-GPU communications (NIXL, NCCL, NVSHMEM), Agile practices

Preferred skills

Open-source contributions to inference frameworks, performance modeling, system-level optimization, mentoring engineers, architectural decision-making

Technologies

vLLM, SGLang, FlashInfer, CUDA, Triton, CUTLASS, NIXL, NCCL, NVSHMEM, Python

Responsibilities

Lead and mentor a high-performing engineering team, drive strategy and roadmap for inference frameworks, partner with compiler and research teams, oversee performance tuning of large-scale models, guide engineers in adopting best practices, represent the team in planning discussions

Seniority

Senior, hands-on IC with management responsibilities

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.