Research Engineer (General)
Core
Build the technical foundation for training and evaluating frontier AI agents by creating environments, improving data quality, and translating real-world workflows into tasks and benchmarks.
Role type
Research Engineer (General)
Builds
Systems for creating, running, evaluating, and improving agent training environments; tools for researchers and data vendors to create higher-quality tasks and feedback loops.
Domain
Artificial Intelligence / Reinforcement Learning / Frontier AI Agents
Deliverable
production ML models | infrastructure
Required skills
Python, Docker, Linux, benchmark design, experiment design, data quality analysis, tool building, pipeline development, independent problem solving
Preferred skills
internal tool building, research infrastructure design, metrics and validation workflow design, competitive programming, Olympiad experience, research publications
Technologies
Python, Docker, Linux
Responsibilities
Build systems for creating, running, evaluating, and improving agent training environments; Design experiments to understand model behavior and data quality issues; Develop tools to help create higher-quality tasks and feedback loops; Work across the full lifecycle of agent training data; Partner with external vendors to improve data engine quality; Build metrics and analyses to validate task and environment usefulness
Seniority
Individual Contributor, early-stage startup

