Member of Technical Staff - Web Crawl Engineer
Core
Build and operate large-scale web crawling infrastructure to continuously discover, acquire, and process content from the internet for frontier AI systems.
Role type
Senior IC distributed systems engineer (web crawling)
Builds
Web-scale data collection infrastructure, distributed crawlers, content extraction pipelines, and dataset delivery systems
Domain
AI infrastructure / Internet-scale data engineering
Deliverable
production ML models
Required skills
Large-scale distributed systems, URL frontier management, crawl orchestration, content extraction, HTML parsing, browser automation, petabyte-scale data processing, systems reliability, observability, performance optimization
Preferred skills
Search engine indexing, anti-bot systems, dynamic web content handling, distributed storage, content deduplication, LLM training data evaluation
Technologies
Ray, Spark, Beam, Flink
Responsibilities
Design and optimize URL discovery and scheduling systems; develop distributed crawlers for diverse web formats; build content extraction and rendering systems; improve crawl coverage and efficiency through experimentation; design recrawling and change detection infrastructure; build observability and monitoring systems; debug production issues and optimize scalability
Seniority
Senior, hands-on IC
