Software Engineer, Pretraining
Core
Build large-scale data systems, pipelines, and crawling infrastructure to transform raw internet-scale data into training-ready datasets for frontier coding models.
Role type
Senior IC software engineer (pretraining data infrastructure)
Builds
High-throughput data pipelines, web crawling systems, and data quality models for frontier model training
Domain
AI/ML pretraining, large-scale distributed systems, data engineering
Deliverable
production ML models | infrastructure
Required skills
large-scale distributed systems, high-throughput pipeline architecture, data quality modeling, web crawling and parsing, system observability, independent debugging, end-to-end ownership
Preferred skills
search infrastructure, AI agent collaboration, rapid iteration, data mixture experimentation
Technologies
distributed systems frameworks, data processing frameworks, web scraping tools, telemetry systems
Responsibilities
Build and own high-throughput data pipelines with end-to-end traceability; Train and ship models for data classification, ranking, and filtering; Design and run scaling-ladder experiments on data quality; Build and scale web crawling systems for document discovery and parsing; Improve URL seeding, scoring, and host scheduling for crawl capacity; Debug and harden complex crawl infrastructure for availability and recovery
Seniority
Senior, hands-on IC
