CareerPlanSign in

Research Engineer, Data Infrastructure (Language Modeling)

*HQ - San Francisco, CA💼 Full-time🗓 2026-09-24 → 2026-09-25

Core

Build performant, scalable data processing infrastructure to acquire, process, and curate massive datasets for pretraining foundational language models.

Role type

Senior Research Engineer, Data Infrastructure

Builds

Scalable data pipelines for ingestion, preprocessing, filtering, deduplication, and augmentation of text datasets for model training.

Domain

Generative AI, Language Modeling, Data Infrastructure

Deliverable

production ML models

Required skills

ML data infrastructure, training data pipelines, dataset versioning, large-scale data loading, data systems design, ablation experiment design, data quality standards, external data vendor management

Preferred skills

Large-scale data processing with Ray/Spark/Kubernetes, pretraining language models

Technologies

Ray, Spark, Kubernetes

Responsibilities

Build and operate performant, scalable data processing infrastructure; Design and run ablation experiments to optimize data mixtures; Partner with research teams on data loading and versioning; Establish rigorous data quality standards; Identify and source novel datasets; Manage relationships and budgets with external data vendors.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.