Senior MLOps Engineer - DSX Enablement
Core
Develop innovative solutions to advance AI infrastructure capabilities, advise experts on ML workload demands, and solve full-stack AI/ML system problems for internal and external customers.
Role type
Senior MLOps Engineer (AI Infrastructure Enablement)
Builds
Custom AI solutions, MLOps pipelines, distributed training/inference systems, and open-source tools on NeoCloud and NVIDIA Cloud Partners platforms.
Domain
AI Infrastructure / High-Performance Computing / Cloud Native
Deliverable
production ML models | infrastructure
Required skills
Linux systems administration, Kubernetes, distributed filesystems, advanced datacenter networking, Python, bash, C++/Go/Rust, ML/DL frameworks, MLOps practices (CI/CD, containerization, GitOps), performance profiling and tuning
Preferred skills
Open-source community contributions, NVIDIA ecosystem (DGX, CUDA, NeMo, RAPIDS, Triton, InfiniBand), security-critical environment experience, deep systems knowledge across hardware to application stack
Technologies
NeoCloud, DGX Cloud, Kubernetes, Linux, Python, C++, Go, Rust, CUDA, NeMo, RAPIDS, Triton, InfiniBand, NVLink, RoCE
Responsibilities
Build and deploy custom AI solutions including distributed training and inference optimization; Act as primary technical contact guiding joint engagements and solving production problems; Profile and tune large-scale workloads to reduce latency and cost; Develop open-source tools and reference architectures for scalable ML systems
Seniority
Senior, hands-on IC