Senior Cloud SRE - AI/ML Platform & GPU Compute
Core
Build and scale reliability foundations for Wayve's AI cloud platform, including the Model Development Platform and large-scale GPU Compute fleets for model training and inference.
Role type
Founding Senior Cloud Site Reliability Engineer (AI/ML Platform)
Builds
Model Development Platform and GPU Compute platform (multi-tenant GPU fleets, scheduling systems)
Domain
Artificial Intelligence / Machine Learning / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes (production clusters), Cloud platforms (AWS/GCP/Azure), Distributed systems, Linux, Scripting (Python/Go/C++), Observability stacks (Datadog/Prometheus/Grafana/OpenTelemetry), Incident management, Capacity planning, Automation/Infrastructure-as-code
Preferred skills
GPU-backed environments, MLOps pipelines, SLO/SLI definition, Early-stage SRE function building
Technologies
Kubernetes, AWS, GCP, Azure, Python, Go, C++, Datadog, Prometheus, Grafana, OpenTelemetry, Terraform
Responsibilities
Own reliability/availability/performance of Model Dev and GPU Compute platforms; Define and operationalize SLOs/SLIs; Lead incident triage and root cause analysis; Design observability systems; Build automation for cluster operations and training workflows; Improve deployment safety via CI/CD hardening.
Seniority
Senior, hands-on IC (Founding role)