Web Scraping Specialist
Core
Build and optimize distributed web crawlers and data pipelines to extract, clean, and store massive amounts of public web data for AI model training.
Role type
Senior IC Web Scraping Engineer
Builds
Distributed web crawlers, data ingestion pipelines, and knowledge graphs
Domain
AI infrastructure / Web-scale data extraction
Deliverable
production ML models
Required skills
Python, JavaScript, asynchronous programming, multithreading, distributed scraping, HTML/CSS/DOM, NoSQL databases (MongoDB, Cassandra), cloud services (AWS, GCP, Azure), data cleaning algorithms
Preferred skills
Machine learning for data categorization, open-source contributions
Technologies
BeautifulSoup, Scrapy, Selenium, MongoDB, Cassandra, AWS, Google Cloud, Azure
Responsibilities
Write and refine code to extract data from complex online sources; handle dynamic content and pagination; clean and format extracted data; manage database storage and integrity; monitor scraping processes and resolve issues
Seniority
Senior, hands-on IC