Software Engineer, Labelling, Data & Automation
Core
Build pipelines, tools, and workflows to extract, label, and curate large-scale real-world data for training AI-first self-driving technology and generative AI simulators.
Role type
Senior IC software engineer (data automation & labelling)
Builds
Scalable data extraction and labelling pipelines, high-reliability systems for data curation, and production deployments of open-set/embedding models.
Domain
Autonomous transportation / Physical AI / Data Engineering
Deliverable
production ML models
Required skills
Python, software engineering fundamentals, large-scale data pipelines, cloud job orchestration, data taxonomy definition, system design, code testing
Preferred skills
ML pipelines (dataset curation, labelling, training), self-driving technology experience, linear algebra and 3D geometry, MapReduce frameworks, open-set/embedding model deployment, infrastructure as code
Technologies
Python, Apache Hadoop, Apache Spark, Apache Airflow, Apache Beam, Google Dataflow, AWS Step Functions, Terraform, CloudFormation
Responsibilities
Design and implement tools and metrics to accelerate autonomy system development; own process and tooling for finding relevant data across petabytes; build high-reliability systems for extracting and labelling data with vendors; define taxonomy and validation rules; manage end-to-end deployment of data solutions; deploy open-set/embedding models to production; champion engineering excellence and code quality; contribute to roadmap planning and delivery.