Research Engineer (Reinforcement Learning)
Core
Build post-training infrastructure for voice and text agents, including environments, synthetic data pipelines, and evaluation systems to ensure reliable model behavior in live conversations.
Role type
Senior IC reinforcement learning research engineer
Builds
Production-ready RL-trained models for voice and text agents
Domain
AI/ML, Reinforcement Learning, Voice AI
Deliverable
production ML models
Required skills
Python, end-to-end model training, synthetic data generation, reward design, GPU management, open-weight model adaptation, evaluation framework design
Preferred skills
GRPO, TRL, verl, OpenRLHF, vLLM, SGLang, FSDA, tool-using agents, execution sandboxes, LoRA, Qwen, Llama
Technologies
Python, GPUs, vLLM, SGLang, FSDP, TRL, verl, OpenRLHF, Qwen, Llama
Responsibilities
Build training environments and verifiers; own the synthetic data pipeline; run end-to-end training experiments; build release evaluation criteria; adapt open-weight base models; ship models to production and iterate based on real usage
Seniority
Senior, hands-on IC