CareerPlanGet AI match score →

Training, Process Management Engineer

London, UK💼 Full-time🗓 2026-03-09 → 2026-07-31

Core

Design, build, and maintain software to orchestrate and monitor machine learning workloads on large supercomputers, ensuring reliability and performance for distributed training runs.

Role type

Senior IC systems engineer (distributed runtime)

Builds

Distributed OS and runtime components for launching, coordinating, and supervising massive training workloads

Domain

AI infrastructure / High-performance computing

Deliverable

production ML models

Required skills

Rust, Python, Linux, asynchronous programming, distributed systems design, performance profiling, systems-level debugging, memory management, fault tolerance design

Preferred skills

Experience developing (not just operating) distributed systems, strong design judgment for ambiguous problems, high-ownership mindset

Technologies

Rust, Python, Linux

Responsibilities

Design and build software to orchestrate and monitor ML workloads on supercomputers; Profile and optimize the software stack for frontier-scale computation; Improve reliability, observability, and fault tolerance for long-running jobs; Debug complex distributed systems issues across large clusters; Respond to evolving needs of ML systems to enable researchers

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗