Infrastructure Engineer (GPU & Compute)
Core
Own image management, system diagnostics, and validation across large-scale bare-metal compute infrastructure with a focus on GPU-enabled systems.
Role type
Senior Infrastructure Engineer (GPU & Compute)
Builds
Automation for provisioning, validation, and system bring-up for AI/ML and HPC workloads
Domain
Cloud infrastructure, GPU compute, HPC
Deliverable
infrastructure
Required skills
Linux systems administration, GPU diagnostics and validation, Python scripting, bare-metal provisioning, hardware debugging, virtualization management
Preferred skills
High-performance interconnects (InfiniBand, NVLink), PXE boot environments, hardware management interfaces (iDRAC, IPMI, Redfish), data center operations, AI/ML workload support
Technologies
NVIDIA DCGM, Linux, Python, PXE, IPMI, Redfish, InfiniBand, NVLink
Responsibilities
Own and evolve systems for image management, deployment, and validation; Run and maintain test clusters for system validation; Diagnose and resolve complex issues across GPUs, drivers, OS, and hardware; Build and maintain automation for provisioning and system bring-up; Manage and operate Linux-based systems in production and validation environments; Partner with Infrastructure, Hardware, and Data Center teams on system bring-up and validation
Seniority
Senior, hands-on IC