CareerPlanGet AI match score →

Principal Machine Learning Infrastructure Engineer

London💼 Full-time🗓 2026-07-07 → 2026-07-31

Core

Design and operate distributed training infrastructure and model serving pipelines for Large Physics Models (neural operators) to enable AI-driven simulation in engineering and manufacturing.

Role type

Principal Machine Learning Infrastructure Engineer

Builds

Distributed training clusters, model serving infrastructure, experiment tracking systems, and deployment pipelines for physics simulation models.

Domain

Deep-tech / AI for Engineering / Numerical Physics / Simulation

Deliverable

production ML models

Required skills

Distributed training (NCCL, FSDP, DDP, pipeline parallelism), Linux systems administration, Kubernetes, SLURM, Python, PyTorch, Cloud GPU infrastructure, Data I/O optimization, CI/CD

Preferred skills

Geometric deep learning, Neural operators, HPC for simulation (CFD/FEA), Model packaging for customer environments, Experiment tracking (Weights & Biases, MLflow), Observability (Prometheus, Grafana)

Technologies

NVIDIA DGX B200, NVLink, InfiniBand, Kubernetes, SLURM, PyTorch, CoreWeave, Prometheus, Grafana, W&B, MLflow

Responsibilities

Design and operate distributed training infrastructure for neural operator architectures; Optimize training pipelines for throughput, fault tolerance, and cost efficiency; Build serving infrastructure for pre-trained LPMs supporting zero-shot inference and uncertainty quantification; Solve data loading bottlenecks for large-scale mesh datasets; Improve developer experience with reliable CI/CD and debugging tools.

Seniority

Principal, hands-on IC with architectural autonomy

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗