CareerPlanSign in

Member of Technical Staff - Web Crawl Engineer

San Francisco, CA💼 Full-time🗓 2026-06-19 → 2026-09-25

Core

Build and operate large-scale web crawling infrastructure to continuously discover, acquire, and process content from the internet for frontier AI systems.

Role type

Senior IC distributed systems engineer (web crawling)

Builds

Web-scale data collection infrastructure, distributed crawlers, content extraction pipelines, and dataset delivery systems

Domain

AI infrastructure / Internet-scale data engineering

Deliverable

production ML models

Required skills

Large-scale distributed systems, URL frontier management, crawl orchestration, content extraction, HTML parsing, browser automation, petabyte-scale data processing, systems reliability, observability, performance optimization

Preferred skills

Search engine indexing, anti-bot systems, dynamic web content handling, distributed storage, content deduplication, LLM training data evaluation

Technologies

Ray, Spark, Beam, Flink

Responsibilities

Design and optimize URL discovery and scheduling systems; develop distributed crawlers for diverse web formats; build content extraction and rendering systems; improve crawl coverage and efficiency through experimentation; design recrawling and change detection infrastructure; build observability and monitoring systems; debug production issues and optimize scalability

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.