CareerPlanSign in

Staff Software Engineer, Inference / Compute Infrastructure Engineering

London💼 Full-time🗓 2026-08-20 → 2026-09-26

Core

Build software state machines that provision, manage, and decommission GPU hardware to create self-service inference clusters for the Research and Inference team.

Role type

Staff Software Engineer, Inference/Compute Infrastructure

Builds

Production-grade infrastructure-as-software platforms, declarative APIs, and control planes for GPU cluster lifecycle management.

Domain

AI Infrastructure, GPU Compute, Cloud Native Systems

Deliverable

production ML models | infrastructure

Required skills

Go, Python, or Rust; Durable workflow orchestration (Temporal, Cadence); State modeling and reconciliation (Kubernetes controllers); Event-driven systems (Kafka, NATS, SQS); Product mindset for internal platforms

Preferred skills

Bare-metal provisioning (PXE, Redfish, BMC); GPU stack knowledge (NCCL, CUDA, InfiniBand); Hyperscaler or GPU cloud experience; Systems programming in Rust or Go

Technologies

Go, Python, Rust, Temporal, Cadence, Kubernetes, Kafka, NATS, SQS, NCCL, CUDA, InfiniBand, RoCE

Responsibilities

Design and implement provisioning state machines for full hardware lifecycle; Build self-service APIs for cluster scaling and teardown; Automate self-healing and node replacement; Ensure pipeline reliability via idempotency and drift detection; Partner with ML teams to encode cluster topology constraints; Engineer infrastructure as typed, tested software with CI/CD

Seniority

Staff, hands-on IC with strategic ownership

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.