Web Scraping Software Engineer
Core
Building and optimizing distributed web crawlers and data pipelines to ingest, segment, and annotate massive volumes of public web data, videos, transcripts, and audio for AI model training.
Role type
Senior IC web scraping software engineer
Builds
Distributed web crawlers, data ingestion pipelines, and annotated datasets for frontier AI labs
Domain
AI infrastructure, web-scale data extraction, distributed systems
Deliverable
production ML models | infrastructure
Required skills
Python, JavaScript, asynchronous programming, multithreading, distributed scraping, HTML/CSS/DOM, NoSQL database design, cloud deployment (AWS/GCP/Azure), data cleaning algorithms
Preferred skills
Machine learning for data categorization, open-source contributions to scraping tools
Technologies
BeautifulSoup, Scrapy, Selenium, MongoDB, Cassandra, AWS, Google Cloud, Azure
Responsibilities
Write and refine code to extract data from complex online sources, handle dynamic content and pagination, clean and format extracted data, manage NoSQL databases for scraped data, monitor scraping processes and resolve issues
Seniority
Senior, hands-on IC
