CareerPlanGet AI match score →

Senior Cloud SRE - AI/ML Platform & GPU Compute

London, United Kingdom💼 Full-time🗓 2026-05-14 → 2026-07-31

Core

Build and scale the reliability foundations of an AI cloud platform, including Model Development and GPU Compute fleets for automated driving systems.

Role type

Founding Senior Cloud Site Reliability Engineer (AI/ML Infrastructure)

Builds

Model Development Platform and large-scale, multi-tenant GPU compute clusters for model training and inference.

Domain

Autonomous driving / Embodied AI / Cloud Infrastructure

Deliverable

production ML models | infrastructure

Required skills

Kubernetes production operations, AWS/GCP/Azure cloud platforms, distributed systems troubleshooting, Linux fundamentals, Python/Go/C++ scripting, observability stack design (Prometheus/Grafana/OpenTelemetry), incident response and root cause analysis, capacity planning for large clusters.

Preferred skills

GPU-backed environment operations, MLOps pipeline experience, Infrastructure-as-Code (Terraform), SLO/SLI definition and reliability program building, founding SRE experience.

Technologies

Kubernetes, AWS, GCP, Azure, Python, Go, C++, Prometheus, Grafana, OpenTelemetry, Terraform

Responsibilities

Own reliability, availability, and performance of Model Dev Platform and GPU Compute; define and operationalize SLOs/SLIs; lead incident triage and post-mortems; design monitoring, logging, and tracing systems; build automation for cluster operations and self-healing patterns; improve CI/CD safety and deployment velocity.

Seniority

Senior, hands-on IC (Founding role)

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗