CareerPlanGet AI match score →

LLM Dataset Engineer

San Francisco💼 Full-time🗓 2026-01-09 → 2026-07-31

Core

Designing taxonomies, filtering heuristics, and post-training pipelines for massive datasets that power foundation LLMs and multimodal models.

Role type

Senior IC LLM Dataset Engineer

Builds

Pre-training corpora, SFT/RLHF/DPO alignment datasets, and multimodal vision/video data pipelines

Domain

AI Infrastructure / Large Language Models / Multimodal AI

Deliverable

production ML models

Required skills

Python (high-performance, multiprocessing, multithreading), petabyte-scale data processing, dataset reconstruction from web crawls, RLHF/DPO pipeline design, data profiling and bias analysis, synthetic data generation

Preferred skills

Computer vision curation, multimodal crawling, taxonomy design for reasoning/coding/math, research background in data-centric AI

Technologies

Spark, Ray, WebDataset, Parquet

Responsibilities

Define mix of web data/code/books/papers for pre-training; implement pipelines for cleaning, deduplication, and signal extraction from petabytes; lead development of SFT and preference modeling datasets; drive acquisition and processing of vision/video data; develop high-throughput ingestion scripts; conduct statistical analysis on training corpora; design pipelines for synthetic data generation

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗