Senior Solutions Architect, NVIDIA Cloud Partner Operations
Core
Hands-on Solutions Architect raising Day 2 operations maturity for NVIDIA Cloud Partners, ensuring health, stability, and efficiency of AI clouds at scale.
Role type
Senior Solutions Architect (Day 2 Operations)
Builds
Operating procedures, reference architectures, automation, and agentic workflows for partner ecosystems
Domain
AI Infrastructure / GPU Cloud / HPC
Deliverable
production ML models | infrastructure
Required skills
Large-scale GPU/HPC/cloud infrastructure operations, Kubernetes/Slurm, Linux, Python/Bash automation, incident/problem management, cost/performance optimization, cross-functional leadership without direct authority
Preferred skills
24/7 operations function design, NVIDIA rack-scale platforms (GB200/GB300 NVL72), Spectrum-X/UFM/Base Command Manager, fleet health benchmarking, GitOps/agent-based remediation
Technologies
DCGM, BMC/Redfish, InfiniBand, NCCL, UFM, Lustre, IBM Storage Scale, WEKA, VAST Data, Kubernetes, Slurm, Prometheus, Grafana, OpenTelemetry, Terraform, Ansible, Argo CD
Responsibilities
Solve hard Day 2 operations problems at scale with partner engineers; prepare operating models for new NVIDIA platforms; improve reliability, performance, and economics via metrics; raise partner Day 2 maturity gaps; convert solutions into ecosystem capabilities; create feedback loops for product/engineering teams
Seniority
Senior, hands-on IC
