Senior Annotation and Data Pipeline Manager
Core
Own the data engine that transforms raw human robot demonstrations into clean, labeled, training-ready datasets to improve the GENE foundation model.
Role type
Senior IC data pipeline and annotation operations manager
Builds
Training datasets for general-purpose robots (Eno) and the underlying foundation model
Domain
Robotics, Embodied AI, Machine Learning
Deliverable
production ML models
Required skills
Python (Pandas, NumPy, Pytorch), SQL, data pipeline architecture, ontology design, vision-language model integration, annotation operation management, metrics tracking (inter-annotator agreement, error rate, throughput)
Preferred skills
Experience scaling data operations at frontier AI labs, hands-on technical leadership in research environments
Technologies
PyTorch, Pandas, NumPy, SQL, vision-language models
Responsibilities
Run the end-to-end data loop from raw trajectory/video to training-ready datasets; Design and maintain the annotation ontology with the model team; Automate annotation using vision-language models for trajectory labeling and data synthesis; Manage internal and vendor labeling operations against quality and delivery schedules; Close the loop by converting robot eval failures into targeted data collection jobs; Track and optimize data quality and throughput metrics
Seniority
Senior, hands-on IC