Research Scientist, Data
Core
Architect and scale data engineering systems to support model training for advanced multimodal foundation models (text, image, audio, video).
Role type
Staff/Lead Research Engineer, Data
Builds
Large-scale data pipelines, robust ML data curation systems, and production-ready data infrastructure for multimodal model training.
Domain
Generative AI, Multimodal Foundation Models, Data Engineering
Deliverable
production ML models
Required skills
Data pipeline architecture, ML data curation, distributed data systems, scalable data ingestion/labeling/filtering/augmentation, data quality assurance, Python, SQL, PySpark, cloud data platforms (AWS/GCP/Azure), privacy and compliance management
Preferred skills
Experience in research or model training environments, expertise with LLMs/VLMs, tool development for dataset management
Technologies
Spark, Hadoop, Ray, Python, SQL, PySpark, AWS, GCP, Azure
Responsibilities
Design and implement large-scale data pipelines for multimodal datasets; curate and manage diverse sensory-rich datasets for pre-training and mid-training; develop tools for data labeling, filtering, and deduplication; optimize data processing for distributed training; ensure data quality, reliability, and ethical compliance; prototype and productionize new dataset management methods.
Seniority
Staff/Lead, hands-on IC with strategic ownership