Manager, Software Engineering - Production AI Inference
Core
Lead production AI inference for NVIDIA Inference Microservices (NIM), enabling customers to deploy optimized enterprise AI models across cloud, data center, and edge environments.
Role type
Senior Engineering Manager (Production AI Inference)
Builds
Production-ready LLM NIMs, optimized inference engines, and validated serving recipes for enterprise customers.
Domain
AI/ML, Large Language Models, Accelerated Computing, Distributed Systems
Deliverable
production ML models
Required skills
Production software delivery, Engineering team management, Process improvement, AI/ML fundamentals, Inference engine optimization, Accelerated computing, Large-scale distributed systems, Security hardening
Preferred skills
Global distributed organization management, Open-source contributions (vLLM, SGLang, TensorRTLLM), Performance optimization (latency/throughput), GPU technologies (CUDA, cuDNN, NCCL), Enterprise/government-ready software delivery
Technologies
CUDA, cuDNN, CUTLASS, cuBLAS, NCCL, NIXL, NVLink, GPUDirect RDMA, Kubernetes, Containers, PyTorch, Triton
Responsibilities
Lead team shipping production-ready LLM NIMs including model onboarding and release readiness; Build predictable operating model via roadmap planning and execution rhythms; Own project execution by anticipating risks and prioritizing timelines; Drive continuous improvement in production workflows; Build and maintain high-performing AI inference engineering team through mentoring and culture building.
Seniority
Senior, hands-on engineering management