Senior Software Engineer, Data Infrastructure
Core
Design and implement scalable, fault-tolerant web crawling and extraction pipelines processing billions of pages.
Role type
Senior Software Engineer (Data Infrastructure)
Builds
Enterprise-scale crawling and extraction platforms for web data acquisition
Domain
B2B Data / Web Crawling / Distributed Systems
Deliverable
production ML models | product features | infrastructure
Required skills
Java, Python, distributed systems, cloud data warehouses (BigQuery/Snowflake), ETL/ELT pipelines, Apache Airflow, Apache Kafka, Kubernetes (GKE/EKS), infrastructure-as-code (Terraform)
Preferred skills
web crawling frameworks (Scrapy), proxy infrastructure, anti-bot evasion, AI/LLM-based extraction, SERP extraction
Technologies
Java, Python, Apache Kafka, GCP (BigQuery, GKE, Vertex AI), Snowflake, Starburst/Trino, Terraform, Scrapy, Kubernetes, Apache Airflow
Responsibilities
Design scalable web crawling and extraction pipelines; Write production-grade code in Java and Python; Build and operate ETL/ELT pipelines; Improve observability and reliability of data systems; Partner with product and data science teams; Contribute to code reviews and documentation
Seniority
Senior, hands-on IC