CareerPlanGet AI match score →

Training, Process Management Engineer

London, England, UK💼 Full-time🗓 2026-07-05 → 2026-07-30

Core

Design, build, and maintain high-performance distributed systems to orchestrate and monitor machine learning workloads on supercomputers, ensuring reliability and observability for research experiments and frontier-scale training runs.

Role type

Senior IC distributed systems engineer (process management)

Builds

Distributed OS and runtime components for launching, coordinating, and supervising massive training workloads

Domain

AI infrastructure / High-performance computing / Distributed systems

Deliverable

production ML models

Required skills

Rust, Python, Linux, asynchronous programming, distributed systems design, performance profiling, systems-level debugging, fault tolerance design

Preferred skills

C++, memory profiling, concurrent systems development, large-scale cluster debugging

Technologies

Rust, Python, Linux

Responsibilities

Design and build software to orchestrate and monitor ML workloads on supercomputers; Profile and optimize the software stack for frontier-scale computation; Improve reliability, observability, and fault tolerance for long-running jobs; Debug complex distributed systems issues across large clusters; Respond to evolving needs of ML systems to enable researchers

Seniority

Senior, hands-on IC

Sourced via linkedin · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on LinkedIn ↗