CareerPlanSign in

Infrastructure Site Reliability Engineer

Gloucestershire💼 Full-time🗓 2026-04-07 → 2026-09-25

Core

Run and evolve a globally scaled, performance-intensive infrastructure stack supporting AI/HPC workloads, ensuring 24/7 stability and security.

Role type

Senior Infrastructure Site Reliability Engineer (IC)

Builds

Resilient, scalable bare-metal and virtualized infrastructure for AI and HPC systems

Domain

AI Infrastructure / High-Performance Computing (HPC)

Deliverable

production ML models | infrastructure

Required skills

Linux administration (Ubuntu), system tuning, disk I/O optimization, hardware-level performance tweaks, Out of Band management (IPMI, Redfish), networking fundamentals, infrastructure scripting (Bash, Python, Ansible), observability (Prometheus, Grafana), orchestration platforms (Kubernetes, MAAS, Tinkerbell), ITSM frameworks

Preferred skills

HPC workloads knowledge, GPU-based infrastructure, InfiniBand networks, HPC performance tuning

Technologies

Ubuntu, IPMI, Redfish, Prometheus, Grafana, Kubernetes, MAAS, Tinkerbell, Bash, Python, Ansible, InfiniBand

Responsibilities

Deploy and operate resilient infrastructure for AI/HPC; optimize Linux system configuration and hardware performance; manage bare-metal infrastructure via IPMI/Redfish; build automation scripts and IaC; maintain observability stack; perform incident postmortems and root cause analysis; mentor junior engineers

Seniority

Senior, hands-on IC with mentorship responsibilities

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.