CareerPlanSign in

Staff Cloud SRE – AI/ML Platform & GPU Compute

London, United Kingdom💼 Full-time🗓 2026-05-06 → 2026-09-25

Core

Founding Staff SRE building reliability foundations for large-scale AI cloud platforms, including Model Development and GPU Compute fleets.

Role type

Staff Cloud Site Reliability Engineer (AI/ML Infrastructure)

Builds

Model Development Platform and multi-tenant GPU compute fleets for model training and inference

Domain

Artificial Intelligence / Machine Learning / Cloud Infrastructure

Deliverable

production ML models | infrastructure

Required skills

SRE/Production Engineer experience, GPU-backed environment operations, MLOps pipeline management, Kubernetes production operations, AWS/GCP/Azure production workloads, distributed systems management, Linux fundamentals, Python/Go/C++ scripting, deep troubleshooting, observability stack design

Preferred skills

Infrastructure-as-code (Terraform), SLO/SLI definition, founding SRE experience, leadership in growing SRE functions

Technologies

Kubernetes, AWS, GCP, Azure, Python, Go, C++, Datadog, Prometheus, Grafana, OpenTelemetry

Responsibilities

Own reliability/availability/performance of Model Dev and GPU platforms, define and operationalize SLOs/SLIs, lead incident response and root cause analysis, design observability systems, build automation for cluster operations and CI/CD

Seniority

Staff, hands-on IC with founding responsibilities

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.