CareerPlanGet AI match score →

Staff Software Engineer, AI Reliability Engineering

London, UK💼 Full-time💰 $325,000–$325,000🗓 2026-04-01 → 2026-07-31

Core

Building reliable, interpretable, and steerable AI systems (Claude) by improving reliability across serving paths, infrastructure, and accelerators.

Role type

Staff Software Engineer, AI Reliability Engineering

Builds

High-availability serving infrastructure, monitoring/observability systems, and incident response capabilities for large language model serving.

Domain

Artificial Intelligence / Large Language Model Serving / Distributed Systems

Deliverable

production ML models | infrastructure

Required skills

distributed systems, infrastructure, reliability engineering, incident response, high-availability system design, cross-team collaboration, holistic system architecture

Preferred skills

SRE or Production Engineer experience, large-scale model serving or training infrastructure operation, ML hardware accelerator expertise (GPUs, TPUs, Trainium), ML-specific networking optimizations (RDMA, InfiniBand), AI-specific observability tools, chaos engineering, open-source infrastructure contributions

Technologies

RDMA, InfiniBand, GPUs, TPUs, Trainium

Responsibilities

Develop Service Level Objectives for LLM serving systems; Design and implement monitoring and observability systems across the token path; Assist in designing high-availability serving infrastructure across regions and cloud providers; Lead incident response for critical AI services; Support reliability of safeguard model serving.

Seniority

Staff, hands-on IC with strategic cross-cutting impact

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗