Software Engineer, Data Ingestion Systems
Core
Build reliable, efficient data ingestion pipelines at petabyte scale to process real-world driving data for annotation, model training, and evaluation.
Role type
Senior IC data engineering engineer (distributed systems)
Builds
High-volume data ingestion pipelines, ETL workflows, and resilient data processing systems for autonomous driving AI
Domain
Autonomous driving / AI Platform / Data Engineering
Deliverable
production ML models
Required skills
Apache Spark, Python, distributed data processing, pipeline debugging, data corruption handling, system optimization, orchestration, failure recovery
Preferred skills
Airflow, Flyte, Databricks, Delta Lake, Scala, Java (via careerplan.io/jobs/424ef55e-7efc-45bd-8f8c-48b8b51cb617-software-engineer-data-ingestion-systems-at-wayve)
Technologies
Apache Spark, Python, Airflow, Flyte, Databricks, Delta Lake, Scala, Java
Responsibilities
Debug and unblock failing ingestion pipelines and processing queues; Investigate corrupt or malformed data from various sources; Optimize Spark jobs and batch-processing pipelines for throughput and efficiency; Improve monitoring, retries, and failure recovery across workflows; Prioritize critical datasets and reduce operational toil
Seniority
Senior, hands-on IC
