Senior AI Data Pipeline Engineer (Autonomous Driving)
Core
Building scalable data processing pipelines and datasets for autonomous driving algorithms, serving millions of scenes to improve ML model development efficiency.
Role type
Senior AI Data Pipeline Engineer (Autonomous Driving)
Builds
Distributed data pipelines, data serving SDKs, and data lakehouses for autonomous driving scene datasets.
Domain
Autonomous driving, mobility AI, big data processing
Deliverable
production ML models
Required skills
Python SDK development, data pipeline job orchestration (Databricks Workflows, Apache Airflow), big data computing engines (Apache Spark), database management (MongoDB, PostgreSQL), data warehouse/lakehouse architecture (Hive, Delta Lake)
Preferred skills
Autonomous vehicle sensor data handling (LiDAR, camera, radar), ML model training lifecycle management, modern AI frameworks (PyTorch, TensorFlow), data governance and privacy implementation
Technologies
Python, Databricks Workflows, Apache Airflow, MongoDB, PostgreSQL, Hive, Delta Lake, Apache Spark
Responsibilities
Develop high-scale data extraction pipelines for raw fleet data; build data labeling pipelines for auto labeling inferences; develop autonomous driving data SDKs; construct data lakehouses for sensor, calibration, and annotation data; optimize data processing and search latency; maintain data platform infrastructure; collaborate with ML and Cloud Infra teams.
Seniority
Senior, hands-on IC