Data Engineer
Core
Design and build data pipelines and flows to support machine learning workflows, from raw data ingestion to model deployment.
Role type
Data Engineer (ML Infrastructure)
Builds
Data pipelines, orchestration flows, and deployment infrastructure for ML models
Domain
Consumer intelligence, retail analytics, machine learning infrastructure
Deliverable
production ML models
Required skills
Python, distributed systems (Dask, Spark), ETL pipeline development, GCP core services (BigQuery, Cloud Storage, GKE), orchestration tools (Airflow, Dagster), Docker, SQL, NoSQL, CI/CD
Preferred skills
ML dataset management, AI model pipeline construction, LangGraph, RAG systems
Technologies
Python, Dask, Spark, GCP (BigQuery, Cloud Storage, Cloud Build, GKE), Airflow, Dagster, Docker, Jupyter, Git, Pandas
Responsibilities
Design and build data pipelines for ML use cases, implement data versioning, deploy and manage flows in orchestration tools, improve and maintain CI/CD pipelines, deploy data scientists' scripts and models to production, implement data quality checks, optimize data processing jobs for performance and cost
Seniority
Mid-level (3–5 years experience), hands-on IC