Engineering Manager, Deep Learning Inference
Core
Lead a world-class engineering team developing and optimizing open-source deep learning inference frameworks (vLLM, SGLang, FlashInfer) for NVIDIA GPUs to enable scalable, real-time AI model deployment.
Role type
Senior Engineering Manager, Deep Learning Inference Software
Builds
Open-source inference frameworks and optimized inference pipelines for LLMs and multimodal generative AI
Domain
AI/ML Infrastructure, GPU Computing, High-Performance Computing
Deliverable
production ML models
Required skills
Technical leadership, C/C++ software design, GPU programming (CUDA, Triton, CUTLASS), performance optimization, multi-GPU communications (NIXL, NCCL, NVSHMEM), Agile practices
Preferred skills
Open-source contributions to inference frameworks, performance modeling, system-level optimization, mentoring engineers, architectural decision-making
Technologies
vLLM, SGLang, FlashInfer, CUDA, Triton, CUTLASS, NIXL, NCCL, NVSHMEM, Python
Responsibilities
Lead and mentor a high-performing engineering team, drive strategy and roadmap for inference frameworks, partner with compiler and research teams, oversee performance tuning of large-scale models, guide engineers in adopting best practices, represent the team in planning discussions
Seniority
Senior, hands-on IC with management responsibilities