HPC Infrastructure Engineer - GPU Clusters
Core
Operate and improve NVIDIA GPU clusters to enable fast, reliable AI model training and research.
Role type
Senior IC HPC Infrastructure Engineer (GPU Clusters)
Builds
Production GPU clusters, automation pipelines, high-performance storage, and job scheduling systems.
Domain
AI/ML Infrastructure, High-Performance Computing, GPU Systems
Deliverable
production ML models | infrastructure
Required skills
Linux server administration, NVIDIA stack (drivers, CUDA, NCCL, DCGM), bare-metal hardware management, high-speed networking (InfiniBand/RoCE), Python/Bash automation, Infrastructure as Code (Ansible/Terraform), performance tuning, storage management.
Preferred skills
ML training workload support, GPU cloud provider evaluation, distributed training failure mode analysis.
Responsibilities
Provision, schedule, monitor, and upgrade GPU fleets; build automated health checks and remediation pipelines; tune job schedulers (Slurm); maintain high-performance storage; diagnose and fix hardware/network performance issues; evaluate rented GPU capacity; perform hands-on hardware racking and cabling.
Seniority
Senior, hands-on IC