Senior HPC Hardware Engineer
Core
Own the full hardware lifecycle for a large-scale HPC and AI compute fleet, ensuring peak performance, availability, and stability for client research and production workloads.
Role type
Senior hands-on IC HPC Hardware Engineer
Builds
Large-scale GPU and CPU node fleets (NVIDIA H200, GB200/NVL72) supporting scientific research, simulations, and analysis
Domain
High-Performance Computing (HPC) and AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
server hardware architecture, bare-metal provisioning, firmware/BIOS lifecycle management, hardware automation (Ansible/Puppet/Chef), Redfish API/BMC/IPMI tooling, GPU/CPU troubleshooting, Linux environments, Python/Bash/PowerShell scripting, capacity planning, security hardening
Preferred skills
OpenStack (Ironic), NVIDIA-SMI/GPU diagnostics, technical leadership/mentoring
Technologies
NVIDIA H200, GB200/NVL72, Ansible, Puppet, Chef, Redfish, iDRAC, iLO, OpenStack Ironic, Linux
Responsibilities
Design and manage high-performance compute fleets; own firmware and BIOS lifecycle; lead troubleshooting of hardware components (CPUs, GPUs, DPUs, NVSwitches, NICs, memory, PSUs, BMCs); automate health checks and issue remediation; validate next-generation AI platforms; collaborate with vendors on hardware issues; perform performance analysis and capacity planning; implement security hardening; mentor junior engineers
Seniority
Senior, hands-on IC with leadership responsibilities