Senior Data Engineer
Core
Design and build production data pipelines on Databricks using Spark Declarative Pipelines (SDP) and PySpark, ensuring data quality, monitoring, and lineage from ingestion to business-ready products.
Role type
Senior IC data engineer (lakehouse)
Builds
Production data pipelines, reusable frameworks, and shared libraries for data products
Domain
Cloud data platforms (Databricks, AWS)
Deliverable
production ML models | product features
Required skills
Spark Declarative Pipelines (SDP/Delta Live Tables), Databricks Asset Bundles (DABs), PySpark, Delta Lake, Unity Catalog, medallion architecture, analytical SQL, Git-based CI/CD, Apache Iceberg
Preferred skills
AWS cloud data platforms, metadata-driven ingestion frameworks, infrastructure-as-code, Snowflake Horizon, Iceberg REST Catalog, enterprise/clinical source system integration (Epic, SAP, Salesforce), legacy ETL migration
Technologies
Databricks, Spark, PySpark, Delta Lake, Unity Catalog, Apache Iceberg, Snowflake, AWS, Kafka, Airflow, dbt
Responsibilities
Design and build production pipelines on Databricks using SDP and PySpark; Define pipelines, jobs, and schedules as code in Databricks Asset Bundles with automated CI/CD; Build data quality, monitoring, and lineage into pipelines; Turn recurring solutions into reusable frameworks and shared libraries; Own pipelines in production regarding performance, cost, reliability, and incident response; Partner with stakeholders to translate requirements into technical designs and mentor engineers
Seniority
Senior, hands-on IC