CareerPlanSign in

Site Reliability Engineer

London, UK🌐 Remote💼 Full-time🗓 2025-08-28 → 2026-08-07

Core

Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support web services and ML workloads, balancing day-to-day operations with long-term software engineering improvements.

Role type

Senior Site Reliability Engineer (SRE)

Builds

Cloud-agnostic platform offering an abstraction layer between science and infrastructure; scalable, highly available, and fault-tolerant infrastructures for web services and ML workloads.

Domain

Cloud computing, High-Performance Computing (HPC), AI/ML infrastructure

Deliverable

production ML models | infrastructure

Required skills

Kubernetes, Terraform, CI/CD, containerization, observability, Python, Go, Bash, networking, security, system administration

Preferred skills

AI/ML environment experience, HPC systems and workload managers (Slurm), modern AI-oriented solutions (Fluidstack, Coreweave, Vast)

Technologies

Kubernetes, Flux, Terraform, Docker, Prometheus, Grafana, ELK Stack, Datadog, CloudFormation, Python, Go, Bash, Slurm, Fluidstack, Coreweave, Vast

Responsibilities

Design and maintain scalable, highly available, and fault-tolerant infrastructures; operate systems and troubleshoot issues in production environments; implement and improve monitoring, alerting, and incident response systems; drive continuous improvement in infrastructure automation, deployment, and orchestration; collaborate with AI/ML researchers to develop solutions for safe and reproducible model-training experiments; document processes and procedures to ensure consistency and knowledge sharing.

Seniority

Senior, hands-on IC

Sourced via adzuna · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.