Staff ML Data Engineer (Datagrid)
Core
Designing and building data systems that power frontier-scale machine learning research and applied AI products, specifically for spatial intelligence and multimodal data.
Role type
Staff ML Data Engineer
Builds
Scalable batch and streaming pipelines for multimodal data (documents, images, spatial metadata) supporting ML training, evaluation, and inference.
Domain
AI & Frontier Models, Spatial Intelligence, Multimodal Data
Deliverable
production ML models
Required skills
SQL, Python, distributed systems, data modeling, dataset lifecycle management, data quality best practices, technical leadership, mentorship
Preferred skills
Large-scale dataset curation, annotation workflows, experiment tracking, reproducibility tooling, lakehouse architectures, event-driven data architectures, infrastructure-as-code, GPU-backed training optimization
Technologies
Databricks, Spark, Kafka, Pub/Sub, Airflow, Dagster, AWS, GCP
Responsibilities
Act as technical lead for data engineering supporting frontier model research; Design and maintain scalable batch and streaming pipelines; Partner with researchers to translate experimental workflows into reliable data systems; Lead development of dataset curation, versioning, and lineage workflows; Establish standards for data quality, validation, and observability; Contribute to data architecture decisions; Identify workflow gaps and run proofs-of-concept; Mentor other engineers through code reviews and design discussions.
Seniority
Staff, hands-on IC with technical leadership