Data Engineer
Core
Build and maintain automated data pipelines on a Databricks lakehouse to transform raw healthcare data into structured, production-ready material for consultants and clients.
Role type
Mid-level Data Engineer
Builds
Automated data pipelines, data models, and production-ready code for healthcare analytics and client engagements
Domain
Healthcare / Data Engineering
Deliverable
production ML models | product features
Required skills
SQL, Python, PySpark, Spark, automated pipeline development, data cleansing, unit/functional/integration testing, Git/GitHub workflow
Preferred skills
Databricks (notebooks, Workflows, Jobs, Delta Lake), Unity Catalog, cloud basics, dbt, dimensional modelling
Technologies
Databricks, PySpark, Spark, Delta Lake, Unity Catalog, dbt, GitHub
Responsibilities
Build and maintain automated data pipelines using PySpark and Spark with monitoring; Develop data quality, validation, and consistency checks; Write unit, functional, and integration tests; Participate in agile ways of working; Build understanding of healthcare data structures and sources; Support business development and bid writing on technical detail
Seniority
Junior to Mid-level, hands-on IC