HPC Infrastructure Site Reliability Engineer
Core
Senior SRE ensuring reliability, performance, and continuous improvement of mission-critical, large-scale GPU-accelerated HPC and AI infrastructure in global data centers.
Role type
Senior Infrastructure Site Reliability Engineer (HPC/AI)
Builds
High-density AI compute platforms, next-generation GPU infrastructure, and global HPC clusters
Domain
High-Performance Computing (HPC) and Artificial Intelligence (AI) Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Linux systems administration and tuning, bare-metal infrastructure management, high-performance networking (InfiniBand, RoCE), GPU ecosystem expertise (NVIDIA), infrastructure automation (IaC, scripting), observability and monitoring, ITIL-aligned incident/problem/change management, root cause analysis, performance benchmarking and validation
Preferred skills
Experience with NVIDIA AI reference architecture, workload schedulers, parallel storage platforms, out-of-band management tooling (IPMI, iLO, iDRAC, Redfish)
Technologies
Linux (Ubuntu), NVIDIA GPUs, InfiniBand, RoCE, Prometheus, Grafana, Bash, Python, Ansible
Responsibilities
Operate and improve high-density AI/HPC infrastructure in a 24/7 production environment; Participate in 24x7x365 on-call rotation for incident response; Troubleshoot complex cross-layer issues in GPU-accelerated environments; Lead performance evaluation and operational acceptance of new HPC infrastructure; Drive continuous service improvement (CSI) through automation and process refinement; Build and maintain infrastructure automation and tooling; Optimise Linux systems (kernel, BIOS/firmware, storage); Configure and operate bare-metal infrastructure; Partner with Platform Engineering and Network Engineering teams; Mentor engineers and act as a technical authority; Feed operational insights into future infrastructure design
Seniority
Senior, hands-on IC with mentorship responsibilities