CareerPlanSign in

Infrastructure Tooling & Observability Engineer( UK)

London💼 Full-time🗓 2026-06-05 → 2026-09-25

Core

Design and build internal control plane tooling and observability systems for a global GPU-as-a-Service infrastructure, transforming high-volume telemetry into actionable insights and automating operational workflows.

Role type

Senior IC Infrastructure Tooling & Observability Engineer

Builds

Internal control plane, observability platforms, automation frameworks, and capacity management systems for a global GPU fleet.

Domain

Cloud Infrastructure / GPU-as-a-Service / HPC

Deliverable

production ML models | product features | infrastructure

Required skills

Ruby (Rails), Go, Ansible, AWX, Kubernetes, Prometheus, Loki, Mimir, Grafana Alloy, REST API design, CI/CD pipelines, SNMP, syslog, capacity planning, incident response

Preferred skills

Kubernetes Certified Administrator, cloud-native observability training, CompTIA+ Security, LPI/LPIC certification

Technologies

Ruby, Go, Ansible, AWX, Kubernetes, Prometheus, Loki, Mimir, Grafana Alloy, GitHub Actions, SNMP, syslog

Responsibilities

Design and evolve internal tooling and observability platforms for large-scale distributed infrastructure; Develop systems to transform high-volume telemetry into actionable insights; Translate SRE reliability requirements into scalable software solutions including automated remediation; Drive automation across infrastructure operations to reduce manual effort; Build tooling for capacity management, performance testing, and benchmarking; Contribute to Continual Service Improvement (CSI) initiatives; Work closely with SRE and Platform Engineering teams to embed observability and reliability.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.