NCX Senior Engineer
Core
Deliver hands-on technical assistance for advanced AI deployments, intricate distributed systems, and ensure customers realize efficient performance from NVIDIA's AI platform across varied environments.
Role type
Senior IC machine-learning infrastructure engineer (cloud/partner support)
Builds
Custom AI solutions on NCP and Neo Cloud platforms, including distributed training, inference optimization, and MLOps pipelines
Domain
AI infrastructure, cloud computing, distributed systems
Deliverable
production ML models
Required skills
Linux systems, distributed computing, Kubernetes, containers, GPU scheduling, Python, Go, PyTorch, TensorFlow, large-scale training/inference support
Preferred skills
NVIDIA ecosystem (DGX, CUDA, NeMo, Triton, NIM, InfiniBand, RoCE), MLOps, cloud-native practices, infrastructure as code
Technologies
NCP, Neo Cloud, DGX Cloud, Kubernetes, PyTorch, TensorFlow, Prometheus, Grafana, OpenTelemetry, Terraform, Ansible
Responsibilities
Build and deploy custom AI solutions on NCP and Neo Cloud platforms; Act as main technical contact for strategic NCPs offering remote and on-site support; Deploy and manage AI workloads across DGX Cloud, NCP data centers, and major CSP environments; Profile and tune large-scale training and inference workloads; Implement and expand NVIDIA reference architectures on partner platforms; Build detailed implementation guides, runbooks, and post-mortem documentation
Seniority
Senior, hands-on IC
