Data Engineer
Core
Building scalable, governed, production-grade data pipelines and analytics products on the Databricks Lakehouse Platform to transform raw data into trusted assets for enterprise analytics and downstream applications.
Role type
Senior Data Engineer (Lakehouse Platform)
Builds
Scalable ETL/ELT pipelines, data products, and analytics infrastructure on Databricks
Domain
Healthcare technology / Data Engineering
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Databricks, Apache Spark (PySpark/Scala), SQL, Python, CI/CD, Delta Lake, Cloud platforms (AWS/Azure/GCP), Data modeling, Data governance
Preferred skills
Medallion design patterns, Data quality frameworks, Regulated environment experience, Healthcare domain knowledge
Technologies
Databricks, Apache Spark, Delta Lake, AWS (EC2, S3, Lambda), Azure DevOps, GitHub Actions, Git, Great Expectations, Deequ
Responsibilities
Design and maintain enterprise data products and pipelines; optimize Spark/Databricks workloads for performance and cost; troubleshoot pipeline issues; collaborate with Data Science and DevOps teams; ensure data privacy compliance (PHI handling); contribute to architecture discussions.
Seniority
Senior, hands-on IC