CareerPlanGet AI match score →

Senior ML Systems Engineer, Frameworks & Tooling

London🌐 Remote💼 Full-time🗓 2025-12-01 → 2026-07-31

Core

Design and maintain the core components of a training framework for frontier-scale language models, enabling fast, reliable, and scalable training on large GPU clusters.

Role type

Senior ML Systems Engineer (Frameworks & Tooling)

Builds

Training framework, distributed training abstractions, monitoring/logging tooling, and reproducible large-scale run infrastructure.

Domain

Large-scale LLM training, distributed systems, HPC infrastructure

Deliverable

production ML models | infrastructure

Required skills

Large-scale distributed training, HPC systems, JAX internals, multi-node cluster orchestration, CUDA/NCCL debugging, containerized environments, performance engineering

Preferred skills

LLM training experience, ML framework contributions (PyTorch, JAX, DeepSpeed), evaluation/serving frameworks, data pipeline optimization

Technologies

JAX, Slurm, Ray, Kubernetes, CUDA, NCCL, Docker, Singularity/Apptainer, GB200/300, AMD, H200/100

Responsibilities

Build and own the training framework for large-scale LLM training; Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO); Improve training throughput and stability on multi-node clusters; Develop tooling for monitoring, logging, and debugging; Collaborate with infra teams on cluster and hardware configurations; Investigate and resolve performance bottlenecks across the ML systems stack

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗