CareerPlanSign in

HPC Infrastructure Engineer - GPU Clusters

United States💼 Full-time🗓 2026-09-03 → 2026-09-25

Core

Operate and improve NVIDIA GPU clusters to enable fast, reliable AI model training and research.

Role type

Senior IC HPC Infrastructure Engineer (GPU Clusters)

Builds

Production GPU clusters, automation pipelines, high-performance storage, and job scheduling systems.

Domain

AI/ML Infrastructure, High-Performance Computing, GPU Systems

Deliverable

production ML models | infrastructure

Required skills

Linux server administration, NVIDIA stack (drivers, CUDA, NCCL, DCGM), bare-metal hardware management, high-speed networking (InfiniBand/RoCE), Python/Bash automation, Infrastructure as Code (Ansible/Terraform), performance tuning, storage management.

Preferred skills

ML training workload support, GPU cloud provider evaluation, distributed training failure mode analysis.

Responsibilities

Provision, schedule, monitor, and upgrade GPU fleets; build automated health checks and remediation pipelines; tune job schedulers (Slurm); maintain high-performance storage; diagnose and fix hardware/network performance issues; evaluate rented GPU capacity; perform hands-on hardware racking and cabling.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.