Intermediate Data Engineer
Core
Designing and maintaining ETL/ELT pipelines to transform messy real-world clinical data into structured, analysis-ready insights for researchers and pharma teams.
Role type
Intermediate Data Engineer
Builds
ETL/ELT pipelines, dbt models, and data layers for clinical data analysis
Domain
Healthcare / Clinical Data Intelligence
Deliverable
production ML models | product features
Required skills
SQL (window functions, CTEs, optimization), Python (production-grade), PySpark, dbt, AWS (S3, MWAA, ECS Fargate, EMR, RDS, Bedrock), data profiling, test suite creation (pytest, Great Expectations)
Preferred skills
Healthcare data standards (HIPAA, FHIR, HL7, OMOP), DevOps (CI/CD, Docker, Terraform/CDK), statistical modeling
Technologies
AWS, Snowflake, dbt, Apache Airflow, PySpark, Python, Cursor, Claude Code
Responsibilities
Design, build, and maintain ETL/ELT pipelines; Optimize pipeline performance and reduce latency/cost; Profile raw datasets to identify quality issues; Build and maintain dbt models; Orchestrate workflows using Apache Airflow on AWS MWAA; Process large-scale data using PySpark on EMR; Document pipelines, models, and assumptions
Seniority
Intermediate, hands-on IC