Research Engineer, Synthetic Data
Core
Build end-to-end synthetic data pipelines that convert domain-specific workflows into realistic, structured, and challenging AI-agent training tasks.
Role type
Research Engineer, Synthetic Data
Builds
Scalable task-generation methods and tooling for AI/ML applications
Domain
AI/ML, Synthetic Data Systems
Deliverable
production ML models
Required skills
Python, Linux, Docker, synthetic data quality criteria, evaluation frameworks, edge case identification
Preferred skills
reinforcement learning, agentic AI workflows, LLM post-training pipelines
Technologies
Python, Linux, Docker
Responsibilities
Design scalable task-generation methods and tooling to mutate, validate, and continuously improve synthetic tasks; Analyze model and agent performance and develop metrics for task diversity, realism, learnability, and quality; Build end-to-end synthetic data pipelines that convert domain-specific workflows into realistic, structured, and challenging AI-agent training tasks
Seniority
Mid-level, hands-on IC