Principal Data Engineer (Apache Spark, dbt, Airflow)
Core
Design and evolve end-to-end data platform solutions, including ingestion, storage, transformation, and serving layers, ensuring scalability and fault tolerance.
Role type
Principal Data Engineer (Cloud & Data Platform)
Builds
Large-scale data platforms, data lakes, data mesh architectures, and reusable engineering frameworks.
Domain
Financial markets, Cloud Infrastructure, Data Engineering
Deliverable
production ML models | product features | infrastructure
Required skills
Large-scale data platform design, AWS cloud engineering, Apache Spark, dbt, Airflow, Terraform, distributed systems, batch and streaming data processing, DataOps principles, infrastructure-as-code, automated testing, observability, data privacy and lineage tracking.
Preferred skills
Lead data engineering role experience, data management capabilities, Confluence, JIRA, ServiceNow.
Technologies
AWS, Apache Spark, dbt, Airflow, Terraform, Redshift, Apache Iceberg, JSON, Parquet, Avro, XML
Responsibilities
Lead design of data lake and data mesh architectures; Design and maintain reusable frameworks and internal developer tooling; Engage stakeholders to ensure optimal data solutions; Drive adoption of DataOps principles including CI/CD and infrastructure-as-code; Collaborate with security teams to embed data privacy and regulatory controls; Support embedded platform data engineers and manage deliveries.
Seniority
Principal, hands-on IC with strategic influence