CareerPlanSign in

Research Engineer (Reinforcement Learning)

North America🌐 Remote💼 Full-time🗓 2026-08-18 → 2026-09-26

Core

Build post-training infrastructure for voice and text agents, including environments, synthetic data pipelines, and evaluation systems to ensure reliable model behavior in live conversations.

Role type

Senior IC reinforcement learning research engineer

Builds

Production-ready RL-trained models for voice and text agents

Domain

AI/ML, Reinforcement Learning, Voice AI

Deliverable

production ML models

Required skills

Python, end-to-end model training, synthetic data generation, reward design, GPU management, open-weight model adaptation, evaluation framework design

Preferred skills

GRPO, TRL, verl, OpenRLHF, vLLM, SGLang, FSDA, tool-using agents, execution sandboxes, LoRA, Qwen, Llama

Technologies

Python, GPUs, vLLM, SGLang, FSDP, TRL, verl, OpenRLHF, Qwen, Llama

Responsibilities

Build training environments and verifiers; own the synthetic data pipeline; run end-to-end training experiments; build release evaluation criteria; adapt open-weight base models; ship models to production and iterate based on real usage

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.