Embedded Data Engineer - ML
Core
Design and build scalable data pipelines, data models, and feature stores to power analytics and machine learning workloads for the travel industry.
Role type
Embedded Data Engineer (ML)
Builds
Cloud-native data applications, reliable datasets for ML use cases, and production data pipelines.
Domain
Travel / Transportation (Rail & Coach) / Machine Learning
Deliverable
production ML models
Required skills
Python, SQL, data pipeline development, feature engineering, cloud data modeling, Spark, Airflow, real-time and batch data processing
Preferred skills
Ray, Parquet, Iceberg, Terraform, Docker, CI/CD pipeline maintenance
Technologies
AWS, DBT, Spark, Airflow, Ray, Parquet, Iceberg, Terraform, Docker, Jenkins, GitHub Actions
Responsibilities
Design and build scalable data pipelines and feature stores; Deploy and maintain cloud-native data applications on AWS; Maintain technical quality and reliability of production data pipelines; Collaborate with ML Engineers and Data Scientists to build datasets for ML use cases; Work with the wider Data Engineering and analytics community to align on best practices.
Seniority
Mid-level to Senior, hands-on IC