CareerPlanSign in

Senior Site Reliability Engineer

Reading, England, United Kingdom💼 Full-time🗓 2026-09-01 → 2026-09-25

Core

Keep production systems running smoothly by fusing engineering principles, operational knowledge, security, and automation to ensure platform/service production excellence for AI and ML workloads.

Role type

Senior Site Reliability Engineer (Infrastructure & AI Platform)

Builds

AI Platform Core services, infrastructure for public cloud deployment, and software delivery lifecycle tooling for ML/LLM features.

Domain

Cloud Infrastructure, DevOps, Machine Learning Operations (MLOps), Large Language Model Operations (LLMOps)

Deliverable

production ML models | infrastructure

Required skills

Kubernetes, Infrastructure as Code (Terraform/CloudFormation), Python or Go, Observability (metrics/logging/tracing), Incident response, Cost governance, Security compliance

Preferred skills

GPU scheduling, Model serving (KServe/RayServe/Triton/vLLM), MLOps pipelines (Kubeflow/MLflow/Feast), LLMOps (prompt management/evals/guardrails), Cross-functional collaboration

Technologies

Kubernetes, Helm, Terraform, AWS, Python, Go, Grafana, Istio, GitHub Actions, GitOps, KServe, RayServe, Triton, vLLM, Kubeflow, MLflow, Feast, W&B

Responsibilities

Manage reliability of ML/LLM workloads including model serving, inference infrastructure, autoscaling, and SLOs; Implement observability for ML models including drift monitoring, tracing, evals, and guardrails; Shape company-wide technical direction and build reusable developer tooling and automation; Package reusable components for open-source tools and ML infrastructure; Enforce secure-by-default infrastructure with compliance audits and cost governance.

Seniority

Senior, hands-on IC

Sourced via workable · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.