Research Engineer - Web Crawlers
Core
Building and operating large-scale, distributed web crawlers to source world-class data from the open web for frontier AI models.
Role type
Senior IC research engineer (web crawling infrastructure)
Builds
Distributed crawling systems, data extraction pipelines, and tooling for monitoring web data quality
Domain
AI/ML data infrastructure, web scale distributed systems
Deliverable
production ML models
Required skills
distributed systems at scale, web crawling/scraping, content extraction from HTML, deduplication at scale, rate-limit handling, data quality evaluation
Preferred skills
Kubernetes, queue-based architectures, multilingual content processing
Technologies
Kubernetes, custom pipelines
Responsibilities
Build and operate large-scale distributed web crawlers, solve hard crawling problems (content extraction, deduplication, freshness strategies), design targeted crawling pipelines for high-value data, create tooling for researchers to request and monitor crawled data
