Data Engineer, Web Scraping
Core
Design, implement, and optimize end-to-end data pipelines for scraping and processing structured and unstructured data to support research and intelligence initiatives.
Role type
Data Engineer (Web Scraping)
Builds
Data pipelines, APIs, dashboards, and data dumps for safety and threat-intelligence teams
Domain
AI safety, threat intelligence, web data collection
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Python, SQL, web scraping/crawling (Beautiful Soup, Selenium, Scrapy), Google Cloud Platform (GCP), data pipeline orchestration, data cleaning/transformation
Preferred skills
Graduate degree in CS/Engineering/Data Science, experience in fast-moving AI research or security environments
Technologies
Google Cloud Platform (Cloud Storage, CloudSQL, Cloud Spanner, Cloud Composer/Airflow, Cloud Run, Pub/Sub), Python, SQL, Beautiful Soup, Selenium, Scrapy
Responsibilities
Design and optimize end-to-end data pipelines for scraping; conduct ad hoc web scraping for research; prepare data for analysis (cleaning, transformation, anonymization); contribute to API development; collaborate with ML engineers and developers to deliver tools and insights
Seniority
Mid-level, hands-on IC