Senior Solutions Architect, First Time Deployment Validation - NVIS
Core
Drive validation of NVIDIA AI factories from first rack power-on through customer handoff, ensuring AI/LLM workloads perform and scale correctly on Linux-based GPU clusters.
Role type
Senior Solutions Architect (AI Infrastructure Validation)
Builds
Validated AI factory environments and launch-ready GPU clusters for NVIDIA's external product launches.
Domain
AI/ML Infrastructure, High-Performance Computing (HPC), Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Linux-based system management, multi-GPU/multi-node cluster operations, NCCL and collective communication patterns (AllReduce, AllToAll), AI/ML benchmarking, Python scripting, Shell/Bash automation, observability (metrics, logs, traces), distributed training frameworks (PyTorch, TensorFlow), performance optimization.
Preferred skills
AI factory or large-scale AI infrastructure build experience, HPC performance engineering background, SRE experience, observability stack expertise, CI-style pipeline development.
Responsibilities
Set up and verify AI factory environments across multi-GPU clusters; execute and analyze AI/LLM benchmarks; investigate and resolve issues with training jobs or benchmarks; build observability solutions for workload behavior; develop automation for benchmarking and regression checks; recommend configuration changes to improve throughput and scaling; collaborate with cross-functional teams to prepare AI factories for customer use.
Seniority
Senior, hands-on IC