Research Engineer / Scientist, Post-training & Reinforcement Learning - London
Core
Develop and train advanced LLMs and VLMs, focusing on training methods for instruction following, tool use, and agentic AI capabilities.
Role type
Senior Research Engineer / Scientist (Post-training & Reinforcement Learning)
Builds
Foundational LLMs and VLMs for agentic AI systems
Domain
Artificial Intelligence, Large Language Models, Reinforcement Learning
Deliverable
production ML models
Required skills
Python, Rust, PyTorch, JAX, TensorFlow, distributed training, SFT, DPO, RLHF/RLVR, reward modelling, offline RL, distillation
Preferred skills
Publications in top-tier AI conferences, PhD or MSc in ML/DL/NLP/CV, large-scale distributed training and inference, training for computer use/agentic settings, RL with sparse rewards, multi-domain training and curriculum learning
Technologies
PyTorch, JAX, TensorFlow
Responsibilities
Develop and train advanced LLMs and VLMs including multimodal architectures; Research and implement training methods for enhanced capabilities like instruction following and tool use; Design and optimize data pipelines and training systems for large-scale distributed training; Collaborate with cross-functional teams to integrate models into agentic AI systems; Evaluate model performance and communicate findings to stakeholders
Seniority
Senior, hands-on IC