Software Engineer, Data Acquisition
Core
Building and maintaining scalable distributed systems for web crawling, data ingestion, and search to support AI model training operations.
Role type
Senior IC software engineer (distributed systems & data acquisition)
Builds
Large-scale web crawlers, data ingestion pipelines, and search indexing systems for AI training data
Domain
Artificial Intelligence / Data Infrastructure
Deliverable
production ML models
Required skills
distributed systems, data processing, web crawling, Kubernetes, Infrastructure-as-Code, backend services, key-value databases, system performance analysis
Preferred skills
large web crawlers experience, experimenting with new technologies
Technologies
Kubernetes, Infrastructure-as-Code, key-value databases
Responsibilities
Lead engineering projects in data acquisition including web crawling and data ingestion; Collaborate with Data Processing, Architecture, and Scaling teams; Work with legal team on compliance and data privacy; Develop and deploy highly scalable distributed systems; Architect and implement algorithms for data indexing and search; Build and maintain backend services for data storage; Deploy solutions in Kubernetes and perform system checks; Analyze experiments on data to provide insights
Seniority
Senior, hands-on IC