Research Engineer, Synthetic Data
Core
Build synthetic data pipelines to create realistic, structured training tasks for frontier AI agents across professional and technical domains.
Role type
Research Engineer (Synthetic Data)
Builds
Synthetic data pipelines, task generation systems, and evaluation tooling for AI agents.
Domain
Artificial Intelligence / Machine Learning / Synthetic Data
Deliverable
production ML models
Required skills
Python, Docker, Linux, synthetic data research methods, end-to-end pipeline building, environment/eval/benchmark design
Preferred skills
First-principles reasoning, edge case identification, unstructured problem solving, independent work in fast-paced environments
Technologies
Python, Docker, Linux
Responsibilities
Collaborate with subject-matter experts to create synthetic tasks; Design synthetic task generation methods; Build systems to mutate, validate, and improve synthetic tasks; Analyze model performance on synthetic tasks; Develop metrics for task diversity, realism, and learnability.
Seniority
Individual Contributor, early-stage startup