CareerPlanSign in

Senior Cloud SRE - AI/ML Platform & GPU Compute

London, United Kingdom💼 Full-time🗓 2026-03-11 → 2026-09-26

Core

Build and scale reliability foundations for Wayve's AI cloud platform, including the Model Development Platform and large-scale GPU Compute fleets for model training and inference.

Role type

Founding Senior Cloud Site Reliability Engineer (AI/ML Platform)

Builds

Model Development Platform and GPU Compute platform (multi-tenant GPU fleets, scheduling systems)

Domain

Artificial Intelligence / Machine Learning / Cloud Infrastructure

Deliverable

production ML models | infrastructure

Required skills

Kubernetes (production clusters), Cloud platforms (AWS/GCP/Azure), Distributed systems, Linux, Scripting (Python/Go/C++), Observability stacks (Datadog/Prometheus/Grafana/OpenTelemetry), Incident management, Capacity planning, Automation/Infrastructure-as-code

Preferred skills

GPU-backed environments, MLOps pipelines, SLO/SLI definition, Early-stage SRE function building

Technologies

Kubernetes, AWS, GCP, Azure, Python, Go, C++, Datadog, Prometheus, Grafana, OpenTelemetry, Terraform

Responsibilities

Own reliability/availability/performance of Model Dev and GPU Compute platforms; Define and operationalize SLOs/SLIs; Lead incident triage and root cause analysis; Design observability systems; Build automation for cluster operations and training workflows; Improve deployment safety via CI/CD hardening.

Seniority

Senior, hands-on IC (Founding role)

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.