Senior Solutions Architect, NVIDIA Cloud Partner Operations
Core
Senior Solutions Architect leading Day 2 operations for NVIDIA Cloud Partners, ensuring cluster health, performance, and efficiency post-deployment.
Role type
Senior Solutions Architect (Cloud Operations & Infrastructure)
Builds
Production-ready AI clouds, operating procedures, reference architectures, and automation for NVIDIA Cloud Partners.
Domain
AI Infrastructure, High-Performance Computing (HPC), Cloud Operations
Deliverable
production ML models | infrastructure
Required skills
Large-scale GPU infrastructure operations, Kubernetes/Slurm, Linux, Python/Bash automation, Incident management, Cost optimization, Distributed systems troubleshooting, Infrastructure as Code (Terraform/Ansible)
Preferred skills
NVIDIA rack-scale platform experience (GB200/GB300), Spectrum-X/UFM/Base Command Manager, 24/7 operations function design, Fleet health benchmarking
Technologies
DCGM, BMC/Redfish, InfiniBand, NCCL, UFM, Lustre, IBM Storage Scale, WEKA, VAST Data, Kubernetes, Slurm, Prometheus, Grafana, OpenTelemetry, Terraform, Ansible, Argo CD
Responsibilities
Solve hard Day 2 operations problems at scale with partner engineers; Prepare new NVIDIA platforms for production; Improve reliability, performance, and economics via metrics; Raise partner Day 2 maturity gaps; Convert solutions into ecosystem capabilities; Create feedback loops for product and engineering teams.
Seniority
Senior, hands-on IC