CareerPlanSign in

Machine Learning Engineer

The City, Central London💼 Full-time💰 $208,000–$208,000🗓 2026-05-30 → 2026-08-04

Core

Optimizing large language model inference on HPC clusters and GPU infrastructure.

Role type

Senior Machine Learning Engineer (Inference Optimization)

Builds

Production inference pipelines for large language models

Domain

Cloud Infrastructure / Large Language Models

Deliverable

production ML models

Required skills

Quantization, PEFT, DeepSpeed, ONNX, TensorRT, PyTorch, multi-LoRa, LoRA Exchange, TitanML, HPC cluster management, LLM scaling, GPU data movement

Preferred skills

Experience with OpenAI, Mistral, Claude, LLaMA models

Technologies

PyTorch, ONNX, TensorRT, TitanML, LoRA Exchange

Responsibilities

Implement quantization and optimization techniques for LLMs, manage inference servers, scale GPU usage and data movement, work with multi-node HPC clusters

Seniority

Senior, hands-on IC

Sourced via adzuna · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.