CareerPlanSign in

Software Engineer, Pretraining

San Francisco💼 Full-time🗓 2026-08-24 → 2026-09-25

Core

Build large-scale data systems, pipelines, and crawling infrastructure to transform raw internet-scale data into training-ready datasets for frontier coding models.

Role type

Senior IC software engineer (pretraining data infrastructure)

Builds

High-throughput data pipelines, web crawling systems, and data quality models for frontier model training

Domain

AI/ML pretraining, large-scale distributed systems, data engineering

Deliverable

production ML models | infrastructure

Required skills

large-scale distributed systems, high-throughput pipeline architecture, data quality modeling, web crawling and parsing, system observability, independent debugging, end-to-end ownership

Preferred skills

search infrastructure, AI agent collaboration, rapid iteration, data mixture experimentation

Technologies

distributed systems frameworks, data processing frameworks, web scraping tools, telemetry systems

Responsibilities

Build and own high-throughput data pipelines with end-to-end traceability; Train and ship models for data classification, ranking, and filtering; Design and run scaling-ladder experiments on data quality; Build and scale web crawling systems for document discovery and parsing; Improve URL seeding, scoring, and host scheduling for crawl capacity; Debug and harden complex crawl infrastructure for availability and recovery

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.