Infrastructure Tooling & Observability Engineer( UK)
Core
Design and build internal control plane tooling and observability systems for a global GPU-as-a-Service infrastructure, transforming high-volume telemetry into actionable insights and automating operational workflows.
Role type
Senior IC Infrastructure Tooling & Observability Engineer
Builds
Internal control plane, observability platforms, automation frameworks, and capacity management systems for a global GPU fleet.
Domain
Cloud Infrastructure / GPU-as-a-Service / HPC
Deliverable
production ML models | product features | infrastructure
Required skills
Ruby (Rails), Go, Ansible, AWX, Kubernetes, Prometheus, Loki, Mimir, Grafana Alloy, REST API design, CI/CD pipelines, SNMP, syslog, capacity planning, incident response
Preferred skills
Kubernetes Certified Administrator, cloud-native observability training, CompTIA+ Security, LPI/LPIC certification
Technologies
Ruby, Go, Ansible, AWX, Kubernetes, Prometheus, Loki, Mimir, Grafana Alloy, GitHub Actions, SNMP, syslog
Responsibilities
Design and evolve internal tooling and observability platforms for large-scale distributed infrastructure; Develop systems to transform high-volume telemetry into actionable insights; Translate SRE reliability requirements into scalable software solutions including automated remediation; Drive automation across infrastructure operations to reduce manual effort; Build tooling for capacity management, performance testing, and benchmarking; Contribute to Continual Service Improvement (CSI) initiatives; Work closely with SRE and Platform Engineering teams to embed observability and reliability.
Seniority
Senior, hands-on IC