CareerPlanGet AI match score →

Research Engineer

London, UK💼 Full-time💰 $120,000–$120,000🗓 2026-05-27 → 2026-07-31

Core

Optimize large-scale training and inference workloads on AI infrastructure to improve performance, scalability, and reliability for real-world AI systems.

Role type

Senior IC Research Engineer (ML Systems & Infrastructure)

Builds

Production AI workloads, inference pipelines, model serving systems, and performance-oriented tooling

Domain

AI Infrastructure / ML Systems / High-Performance Computing

Deliverable

production ML models

Required skills

Deep learning frameworks (PyTorch), large-scale training/inference workloads, distributed systems and parallelism strategies, software engineering fundamentals (API design, tooling, debugging), performance bottleneck analysis

Preferred skills

Inference optimization techniques (quantization, speculative decoding, mixed precision), CUDA/Triton/TensorRT/vLLM/SGLang/Dynamo, open-source contributions, startup experience, advanced degree in AI/ML/systems

Technologies

PyTorch, CUDA, Triton, TensorRT, vLLM, SGLang, Dynamo, NVIDIA GPUs, TPUs

Responsibilities

Optimize large-scale training and inference workloads across GPUs and distributed systems; Work with customers to analyze workloads and improve deployed AI system performance; Develop and improve inference pipelines and model serving systems; Design profiling, debugging, and observability tools; Partner with hardware vendors to support diverse compute backends; Contribute to open-source projects

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗