CareerPlanSign in

Platform Site Reliability Engineer

Gloucestershire💼 Full-time🗓 2026-05-06 → 2026-09-25

Core

Design and operate AI-native cloud platforms, managing GPU-native workloads, multi-tenant control planes, and high-performance AI systems.

Role type

Senior Platform Site Reliability Engineer (Kubernetes & Linux)

Builds

AI infrastructure platform, Kubernetes clusters, observability stack, and automation tooling.

Domain

Cloud Infrastructure, AI Systems, High-Performance Computing

Deliverable

production ML models | infrastructure

Required skills

Kubernetes administration, Linux system tuning, infrastructure as code, observability stack management, incident management, networking fundamentals, automation scripting

Preferred skills

AI workload orchestration, HPC supportability, ITSM frameworks, mentorship

Technologies

Kubernetes, Prometheus, Grafana, Bash, Python, Ansible, Linux (Ubuntu), TCP/IP, DNS, DHCP

Responsibilities

Deploy and manage scaled Kubernetes clusters for AI workloads; optimize Linux system configuration and kernel parameters; build automation scripts and IaC for platform lifecycle; maintain observability stack and conduct incident postmortems; mentor junior engineers and consult on operational requirements.

Seniority

Senior, hands-on IC with mentorship responsibilities

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.