CareerPlanGet AI match score →

Research Engineer - Contextual Bandits & RL

London💼 Full-time🗓 2026-06-12 → 2026-07-31

Core

Build decision-making models for in-store hyper-personalization using offline contextual bandits and reinforcement learning, learning from logged human interaction data.

Role type

Research Engineer (Offline Contextual Bandits & RL)

Builds

Decision-making models for single-step and multi-step customer journeys in retail

Domain

Retail / Hyper-personalization / Offline Reinforcement Learning

Deliverable

production ML models

Required skills

Contextual bandits, Reinforcement learning, Counterfactual learning, Transformers, Graph Neural Networks (GNNs), Python, Off-policy evaluation (OPE), Dataset design, Production-level code debugging

Preferred skills

Offline policy learning and evaluation methods (IPS, doubly-robust), Bandit algorithms and exploration strategies, Recommender systems and ranking, Data pipeline construction

Technologies

Python, Transformers, GNNs

Responsibilities

Develop and productionize offline contextual bandit and offline RL methods; Build rigorous off-policy evaluation and counterfactual validation; Formulate single-step and multi-step decision processes based on real retail interactions; Advance representation learning for decision-making; Translate research ideas into robust systems including deployment and monitoring; Collaborate cross-functionally to turn ambiguous product goals into concrete ML objectives

Seniority

Mid-Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗