Staff Data Engineer
Core
Design, build, and operate data infrastructure powering large-scale model training and inference for AI products in the general aviation industry.
Role type
Staff Data Engineer (AI Platform)
Builds
Data pipelines, data lakehouse, feature stores, and data quality mechanisms for ML workloads.
Domain
Aerospace / Artificial Intelligence / Data Engineering
Deliverable
production ML models
Required skills
High-throughput data pipeline design, open table formats (Iceberg, Paimon), columnar formats (Parquet), object storage (S3), high-throughput messaging (Kafka, Pulsar), SQL, StarRocks, AWS data services, containerization (Docker/Kubernetes), orchestration (Airflow/Prefect/Dagster), data quality instrumentation, lineage tracking, anomaly detection, feature store design, dataset versioning.
Preferred skills
CDC replication (Debezium, Airbyte), audio/time-series data preprocessing, aerospace/aviation domain experience.
Responsibilities
Design and maintain fault-tolerant ingestion and transformation pipelines; define table formats and partitioning strategies for the data lakehouse; instrument pipelines with data quality checks and anomaly detection; partner with ML engineers on feature stores and data contracts; optimize query performance and unblock training runs.
Seniority
Staff, hands-on IC with strategic ownership