Data Engineer (Python)
Core
Build and scale health data pipelines enabling data discoverability, linkage, privacy, machine learning, and analytics across the healthcare ecosystem.
Role type
Data Engineer (Python, Spark, Snowflake)
Builds
Foundational data infrastructure supporting regulatory-grade health decisions at scale
Domain
Healthcare / Life Sciences / Data Infrastructure
Deliverable
production ML models
Required skills
Python, Apache Spark, Airflow, Snowflake, agentic code generation tools (Claude Code/Codex)
Preferred skills
dimensional/analytics data modeling, AWS data environments, data governance and quality frameworks, on-premises deployment architecture, application security best practices
Technologies
Python, Spark, Airflow, Snowflake, AWS
Responsibilities
Design and maintain scalable data pipelines; ensure pipeline reliability, data quality, and observability; partner with Science and Product to define data requirements; implement automated testing and optimize processing performance; design and maintain data validation frameworks aligned with regulatory requirements
Seniority
Mid-to-Senior (2+ years for junior, 5+ for senior)