CareerPlanGet AI match score →

Manager, Software Engineering - Production AI Inference

US, CA, Santa Clara💼 Full-time💰 $224,000–$224,000🗓 2026-07-02 → 2026-07-31

Core

Lead production AI inference for NVIDIA Inference Microservices (NIM), enabling customers to deploy optimized enterprise AI models across cloud, data center, and edge environments.

Role type

Senior Engineering Manager (Production AI Inference)

Builds

Production-ready LLM NIMs, optimized inference engines, and validated serving recipes for enterprise customers.

Domain

AI/ML, Large Language Models, Accelerated Computing, Distributed Systems

Deliverable

production ML models

Required skills

Production software delivery, Engineering team management, Process improvement, AI/ML fundamentals, Inference engine optimization, Accelerated computing, Large-scale distributed systems, Security hardening

Preferred skills

Global distributed organization management, Open-source contributions (vLLM, SGLang, TensorRTLLM), Performance optimization (latency/throughput), GPU technologies (CUDA, cuDNN, NCCL), Enterprise/government-ready software delivery

Technologies

CUDA, cuDNN, CUTLASS, cuBLAS, NCCL, NIXL, NVLink, GPUDirect RDMA, Kubernetes, Containers, PyTorch, Triton

Responsibilities

Lead team shipping production-ready LLM NIMs including model onboarding and release readiness; Build predictable operating model via roadmap planning and execution rhythms; Own project execution by anticipating risks and prioritizing timelines; Drive continuous improvement in production workflows; Build and maintain high-performing AI inference engineering team through mentoring and culture building.

Seniority

Senior, hands-on engineering management

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗