Engineering Manager, Deep Learning Inference
Core
Lead a world-class engineering team advancing AI model deployment software, specifically optimizing open-source inference frameworks (SGLang, vLLM, FlashInfer) for NVIDIA GPUs to enable real-time inference from datacenters to edge devices.
Role type
Senior IC engineering manager (deep learning inference)
Builds
Open-source inference frameworks and optimized inference pipelines for LLMs and generative AI
Domain
AI/ML infrastructure, GPU computing, high-performance computing
Deliverable
production ML models
Required skills
C/C++ software design, GPU programming (CUDA, Triton, CUTLASS), performance tuning and profiling, multi-GPU communications (NIXL, NCCL, NVSHMEM), team leadership and mentorship, open-source framework development
Preferred skills
Python proficiency, experience with PyTorch/TensorRT-LLM, publications/patents on LLM serving, expertise in distributed inference architectures
Technologies
CUDA, Triton, CUTLASS, NIXL, NCCL, NVSHMEM, SGLang, vLLM, FlashInfer, PyTorch, TensorRT-LLM
Responsibilities
Lead and scale a high-performing engineering team, guide strategy and roadmap for OSS inference frameworks, partner with compiler and research teams for end-to-end optimization, oversee performance tuning of large-scale models, guide engineers in adopting best practices for GPU programming, represent the team in roadmap planning
Seniority
Senior, hands-on IC with management responsibilities