Data Engineer, Application Software
Core
Design and deliver scalable data pipelines for model-based autonomous driving, transforming diverse data into structured, model-ready datasets for ML workflows.
Role type
Data Engineer
Builds
Scalable data pipelines, data quality systems, and curated datasets for autonomous driving models
Domain
Autonomous driving, Machine Learning, Robotics
Deliverable
production ML models
Required skills
scalable data pipeline development, Python, SQL, PySpark, distributed data processing, data quality assurance, workflow orchestration, autonomous driving data concepts (sensor characteristics, coordinate transformations, calibration, GNSS/IMU)
Preferred skills
calibration, perception, imitation learning, trajectory prediction, third-party dataset ingestion, automated driving scenario taxonomies
Technologies
Python, SQL, PySpark, Airflow, Flyte, Ray
Responsibilities
build and improve scalable data pipelines for ML workflows, ingest and transform large-scale real-world and synthetic datasets, develop data quality checks and monitoring, curate data for scenario diversity, improve pipeline performance and reliability, collaborate with ML engineers and Data Corpus teams
Seniority
Mid-to-Senior, hands-on IC