Principal Data Engineer (AWS, Databricks, Ai, ML Flow, Data Architecture, Apache Airflow)
Core
Design and build scalable data foundations powering advanced analytics, machine learning, generative AI, and agentic AI solutions.
Role type
Principal Data Engineer (hands-on IC with technical leadership)
Builds
Scalable batch, streaming, and event-driven data platforms; governed lakehouse architectures; reusable data products and APIs.
Domain
Financial services / Cloud-native data engineering / AI/ML infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Python, PySpark, SQL, Apache Spark, AWS cloud platforms, data lakehouse architecture, Apache Airflow, CI/CD, Infrastructure as Code, data governance, observability
Preferred skills
MLflow, Azure Machine Learning, Unity Catalog, Kafka, Docker, Kubernetes, Terraform, responsible AI practices
Technologies
AWS, Databricks, Apache Airflow, PySpark, Python, SQL, Apache Spark, MLflow, Kafka, Docker, Kubernetes, Terraform
Responsibilities
Lead architecture and engineering of scalable data platforms; build data ingestion and transformation pipelines; enable full AI/ML lifecycle from feature engineering to inference; define engineering standards and governance; provide hands-on technical leadership and mentorship.
Seniority
Principal, hands-on IC with strategic influence