Lead Research Engineer, Data Quality
Core
Lead the data quality team in building systems that evaluate thousands of tasks across RL environments, synthetic data pipelines, benchmarks, and domain-specific workflows for frontier AI agents.
Role type
Senior individual-contributor and team-lead for AI evaluation and data quality
Builds
QC systems, evals, benchmarks, synthetic data pipelines, and model evaluation infrastructure
Domain
Artificial Intelligence / Reinforcement Learning / Data Quality Engineering
Deliverable
production ML models
Required skills
Python, Docker, Linux, building QC systems, AI evals, post-training, designing metrics and experiments, task mutation checks, trajectory auditing, failure-mode analysis
Preferred skills
leading technical teams, translating qualitative research insights into production systems, mentoring engineers, operating in early-stage startup environments
Technologies
Python, Docker, Linux
Responsibilities
Lead the data quality team in building systems that evaluate thousands of tasks across RL environments, synthetic data pipelines, benchmarks, and domain-specific workflows; Define the data quality strategy by building QC systems, enforcing standards, and designing experiments to grade agent outputs; Develop and implement methods for validating synthetic data at scale, including failure-mode analysis, task mutation checks, and trajectory auditing; Partner with research engineers, domain experts, and data vendors to diagnose quality issues and improve data generation workflows; Translate qualitative research insights into production systems: internal tools, dashboards, validation pipelines, and feedback loops; Mentor other research engineers to maintain a high bar for technical rigor, clarity, and execution speed
Seniority
Senior, hands-on IC and team lead