CareerPlanGet AI match score →

Member of Engineering (Scalability)

UK🌐 Remote💼 Full-time🗓 2025-11-08 → 2026-07-31

Core

Building distributed training and inference infrastructure for Large Language Models (LLMs) to enable source code generation.

Role type

Senior IC distributed systems engineer (LLM infrastructure)

Builds

High-reliability distributed training and inference systems for foundational LLMs

Domain

AI infrastructure / Large Language Models / High-performance computing

Deliverable

production ML models

Required skills

Linux kernel debugging, distributed systems, fault tolerance, NCCL, PyTorch, C/C++, CUDA API, algorithmic design

Preferred skills

GPU architecture knowledge, checkpointing optimization, observability, K8s stack

Technologies

PyTorch, CUDA, NCCL, Linux kernel, C/C++, Python, Cython, K8s

Responsibilities

Troubleshoot hardware problems during large-scale training, minimize GPU idle time during faults, design tools for training recovery, improve checkpointing performance and reliability, write high-performance code in Python/C/C++/CUDA

Seniority

Senior, hands-on IC

Sourced via adzuna · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Adzuna ↗