Machine Learning Engineer - Data Pipeline
Core
Build end-to-end data pipelines (collection, extraction, filtering, synthetic generation) to power domain-specific AI model training.
Role type
Senior IC machine learning engineer (data infrastructure)
Builds
Large-scale distributed data processing systems and pipelines for AI training data
Domain
Generative AI, data engineering, distributed systems
Deliverable
production ML models | infrastructure
Required skills
deep learning frameworks (PyTorch), large-scale distributed data processing, Python, data pipeline architecture, data quality improvement algorithms
Preferred skills
web crawling tools (Scrapy, Selenium), data processing frameworks (Hadoop, Datasketch), building bespoke data libraries, multilingual data handling
Technologies
PyTorch, Ray, Docker, Kubernetes, Scrapy, Selenium, Hadoop, Datasketch, key-value databases
Responsibilities
Design and develop data processing pipelines for extraction, filtering, and labeling; Implement ML models to improve data quality and diversity; Lead engineering projects in data acquisition (web crawling, ingestion); Develop and deploy scalable distributed systems for terabytes of data; Architect algorithms for data indexing and search; Build and maintain backend data storage services; Deploy solutions in Kubernetes IaC environments
Seniority
Senior, hands-on IC