CareerPlanSign in

Research Engineer / Scientist, Post-training & Reinforcement Learning - London

Hybrid London💼 Full-time🗓 2026-04-14 → 2026-09-26

Core

Develop and train advanced LLMs and VLMs, focusing on training methods for instruction following, tool use, and agentic AI capabilities.

Role type

Senior Research Engineer / Scientist (Post-training & Reinforcement Learning)

Builds

Foundational LLMs and VLMs for agentic AI systems

Domain

Artificial Intelligence, Large Language Models, Reinforcement Learning

Deliverable

production ML models

Required skills

Python, Rust, PyTorch, JAX, TensorFlow, distributed training, SFT, DPO, RLHF/RLVR, reward modelling, offline RL, distillation

Preferred skills

Publications in top-tier AI conferences, PhD or MSc in ML/DL/NLP/CV, large-scale distributed training and inference, training for computer use/agentic settings, RL with sparse rewards, multi-domain training and curriculum learning

Technologies

PyTorch, JAX, TensorFlow

Responsibilities

Develop and train advanced LLMs and VLMs including multimodal architectures; Research and implement training methods for enhanced capabilities like instruction following and tool use; Design and optimize data pipelines and training systems for large-scale distributed training; Collaborate with cross-functional teams to integrate models into agentic AI systems; Evaluate model performance and communicate findings to stakeholders

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.