CareerPlanSign in

Machine Learning Systems & Infrastructure Engineer

London💼 Full-time🗓 2026-05-06 → 2026-09-26

Core

Build and own scalable ML systems to train, evaluate, and serve large diffusion-based generative world models for robotics, AR/VR, gaming, and cinema.

Role type

Senior IC machine learning systems & infrastructure engineer

Builds

Scalable training stacks, data ingestion pipelines, experiment orchestration, and model serving endpoints for large foundation models

Domain

Generative AI, computer vision, 3D world modeling, cloud infrastructure

Deliverable

production ML models | infrastructure

Required skills

Python, PyTorch (DDP/FSDP), distributed training debugging, data pipeline engineering, GPU compute optimization, Kubernetes, Terraform, SQL, observability tooling

Preferred skills

Experience with large-scale object storage, ML workflow orchestration, CI/CD for ML workflows

Technologies

PyTorch, Docker, Kubernetes, Terraform, Prometheus, Grafana, OpenTelemetry, Kubeflow Pipelines, Airflow, Volcano, Slurm, MLflow, Weights & Biases, Modal, Triton, Playwright, AWS/GCP/Azure

Responsibilities

Design and operate scalable training stacks and data ingestion pipelines; optimize distributed training performance and stability; manage containerization, IaC, and CI/CD pipelines; implement observability, monitoring, and incident response; collaborate with researchers to unblock training and data work

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.