Compute Platform Engineer
Core
Designing, configuring, and managing high-performance compute infrastructure (CPU/GPU nodes) to support critical research and production workloads.
Role type
Senior IC infrastructure engineer (HPC/AI hardware)
Builds
High-performance compute platforms for scientific research and production
Domain
High-performance computing (HPC) and AI infrastructure
Deliverable
infrastructure
Required skills
HPE server infrastructure management, NVIDIA GPU management, server architecture (UEFI/BIOS, PCIe), out-of-band management (iLO, BMC), hardware troubleshooting, automation (Ansible, Terraform, CI/CD), Linux in high-performance environments, network concepts (DNS, DHCP, VLANs, switching, routing), capacity planning, Infrastructure as Code (IaC)
Preferred skills
Kubernetes, Openstack
Responsibilities
Designing and managing HPC infrastructure with GPU and CPU nodes, managing firmware/BIOS lifecycle, troubleshooting hardware components, automating health checks and onboarding workflows, monitoring hardware performance, collaborating with vendors on firmware issues, implementing security hardening, mentoring junior engineers, performing diagnostics and capacity planning
Seniority
Senior, hands-on IC