Senior HPC Platform Architect
Core
Architect and optimize large-scale HPC clusters and data center infrastructure powering NVIDIA's silicon design and AI workloads.
Role type
Senior HPC Platform Architect
Builds
High-performance compute clusters, storage systems, and networking fabrics for data centers
Domain
High-Performance Computing (HPC), Data Center Infrastructure, Systems Engineering
Deliverable
production ML models | infrastructure
Required skills
HPC infrastructure architecture, data center design, Linux kernel tuning, workload management (LSF/Slurm), performance benchmarking, Python/Bash scripting, rack layout planning, network topology design
Preferred skills
GPU cluster optimization, multi-site topology design, capacity planning, observability framework development
Technologies
InfiniBand, Ethernet, NVLink, LSF, Slurm, NVMe, NUMA, huge pages
Responsibilities
Own data center architecture reviews for new HPC clusters; Lead performance benchmarking and profiling of cluster infrastructure; Drive infrastructure optimization at scheduler, hardware, and OS/kernel levels; Evaluate new hardware, storage systems, and networking fabrics; Collaborate on cluster health and capacity planning; Improve infrastructure observability and documentation
Seniority
Senior, hands-on IC
