Research Engineer, Data Infrastructure (Language Modeling)
Core
Build performant, scalable data processing infrastructure to acquire, process, and curate massive datasets for pretraining foundational language models.
Role type
Senior Research Engineer, Data Infrastructure
Builds
Scalable data pipelines for ingestion, preprocessing, filtering, deduplication, and augmentation of text datasets for model training.
Domain
Generative AI, Language Modeling, Data Infrastructure
Deliverable
production ML models
Required skills
ML data infrastructure, training data pipelines, dataset versioning, large-scale data loading, data systems design, ablation experiment design, data quality standards, external data vendor management
Preferred skills
Large-scale data processing with Ray/Spark/Kubernetes, pretraining language models
Technologies
Ray, Spark, Kubernetes
Responsibilities
Build and operate performant, scalable data processing infrastructure; Design and run ablation experiments to optimize data mixtures; Partner with research teams on data loading and versioning; Establish rigorous data quality standards; Identify and source novel datasets; Manage relationships and budgets with external data vendors.
Seniority
Senior, hands-on IC