CareerPlanSign in

Engineering Manager, Deep Learning Inference

6 Locations💼 Full-time💰 $224,000–$224,000🗓 2026-07-28 → 2026-09-26

Core

Lead a world-class engineering team advancing AI model deployment software, specifically optimizing open-source inference frameworks (SGLang, vLLM, FlashInfer) for NVIDIA GPUs to enable real-time inference from datacenters to edge devices.

Role type

Senior IC engineering manager (deep learning inference)

Builds

Open-source inference frameworks and optimized inference pipelines for LLMs and generative AI

Domain

AI/ML infrastructure, GPU computing, high-performance computing

Deliverable

production ML models

Required skills

C/C++ software design, GPU programming (CUDA, Triton, CUTLASS), performance tuning and profiling, multi-GPU communications (NIXL, NCCL, NVSHMEM), team leadership and mentorship, open-source framework development

Preferred skills

Python proficiency, experience with PyTorch/TensorRT-LLM, publications/patents on LLM serving, expertise in distributed inference architectures

Technologies

CUDA, Triton, CUTLASS, NIXL, NCCL, NVSHMEM, SGLang, vLLM, FlashInfer, PyTorch, TensorRT-LLM

Responsibilities

Lead and scale a high-performing engineering team, guide strategy and roadmap for OSS inference frameworks, partner with compiler and research teams for end-to-end optimization, oversee performance tuning of large-scale models, guide engineers in adopting best practices for GPU programming, represent the team in roadmap planning

Seniority

Senior, hands-on IC with management responsibilities

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.