Research Crawling Engineer
Core
Design and operate large-scale web data acquisition systems for research and model development, spanning distributed systems, scraping infrastructure, and data pipelines.
Role type
Senior IC Research Crawling Engineer
Builds
Distributed crawlers, high-throughput data collection systems, and datasets for research and model training
Domain
Web-scale data infrastructure, AI/ML data pipelines, distributed systems
Deliverable
production ML models
Required skills
Go, Rust, Python, Java, C++, web crawler development, large-scale data pipelines, HTTP, networking, browser behavior, distributed systems, parallel processing, large dataset handling (TB–PB scale), adversarial environment debugging
Preferred skills
NLP pipelines, dataset curation for ML, LLM pretraining data, retrieval systems, headless browsers (CDP, Playwright, Puppeteer), proxy systems, IP rotation, request orchestration, data quality evaluation, cloud/bare-metal infrastructure
Technologies
Go, Rust, Python, Java, C++, Chrome DevTools Protocol, Playwright, Puppeteer
Responsibilities
Build and maintain large-scale web crawlers across diverse domains; Design high-throughput, fault-tolerant systems for data collection; Handle anti-bot systems, rate limits, and dynamic/JS-heavy sites; Develop pipelines for cleaning, deduplication, filtering, and normalization; Construct and maintain datasets for research and model training; Monitor crawl performance, coverage, and data quality; Collaborate with research teams to align data collection with modeling needs; Optimize infrastructure for cost, latency, and reliability
Seniority
Senior, hands-on IC
