Data Engineer
Core
Build and maintain distributed data pipelines and web scraping infrastructure to collect, process, and deliver massive datasets for training AI models.
Role type
Senior Data Engineer (Web Scale / AI Data)
Builds
Distributed crawlers, data ingestion pipelines, and large-scale datasets for frontier AI labs
Domain
AI/ML Data Infrastructure, Web Scraping, Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Python (advanced, async, multiprocessing), Web scraping at scale, Distributed data pipelines, Data warehousing, Docker & Kubernetes, Linux & bare-metal ops, CI/CD
Preferred skills
Databend, ClickHouse, BigQuery, Helm charts, ArgoCD
Technologies
Python, Celery, Kafka, RabbitMQ, Databend, ClickHouse, BigQuery, Docker, Kubernetes, Linux, GitHub Actions, ArgoCD
Responsibilities
Maintain and optimize database queries and data systems; Create and improve data pipelines for dataset collection and processing; Develop and maintain web scraping tools and scripts; Monitor and troubleshoot pipeline issues to ensure data quality; Document engineering workflows and technical decisions; Participate in R&D projects to improve data products.
Seniority
Senior, hands-on IC