CareerPlanGet AI match score →
💼 Full-time🗓 2026-06-25

Core

Designing and maintaining ETL/ELT pipelines to transform messy real-world clinical data into structured, analysis-ready insights for researchers and pharma teams.

Role type

Intermediate Data Engineer

Builds

ETL/ELT pipelines, dbt models, and data layers for clinical data analysis

Domain

Healthcare / Clinical Data Intelligence

Deliverable

production ML models | product features

Required skills

SQL (window functions, CTEs, optimization), Python (production-grade), PySpark, dbt, AWS (S3, MWAA, ECS Fargate, EMR, RDS, Bedrock), data profiling, test suite creation (pytest, Great Expectations)

Preferred skills

Healthcare data standards (HIPAA, FHIR, HL7, OMOP), DevOps (CI/CD, Docker, Terraform/CDK), statistical modeling

Technologies

AWS, Snowflake, dbt, Apache Airflow, PySpark, Python, Cursor, Claude Code

Responsibilities

Design, build, and maintain ETL/ELT pipelines; Optimize pipeline performance and reduce latency/cost; Profile raw datasets to identify quality issues; Build and maintain dbt models; Orchestrate workflows using Apache Airflow on AWS MWAA; Process large-scale data using PySpark on EMR; Document pipelines, models, and assumptions

Seniority

Intermediate, hands-on IC

Rewrite
## About the Role Century Health is a clinical data intelligence company turning messy real-world clinical data into structured, analysis-ready insights for researchers and pharma teams. We're looking for an Intermediate Data Engineer (3–5 years) to be a core builder on our data platform — designing ETL pipelines, profiling raw clinical datasets, and ensuring data flowing through our systems is clean and reliable. ## What You'll Do - Design, build, and maintain ETL/ELT pipelines ingesting data from CSVs, Parquet, XLSX, APIs, and databases - Optimize pipeline performance — tune queries, manage compute, reduce latency and cost - Profile raw datasets to identify quality issues: missing values, duplicates, schema drift, outliers - Build and maintain dbt models for clean, documented, analysis-ready data layers - Orchestrate workflows using Apache Airflow on AWS MWAA - Process large-scale data using PySpark on EMR - Collaborate with ML engineers and GTM teams on downstream use cases - Document pipelines, models, assumptions, and known issues clearly ## What We're Looking For ### Must-Have - 3–5 years of professional data engineering experience - Strong SQL — window functions, CTEs, query optimization - Solid Python — clean, modular, production-grade code - Hands-on PySpark for large-scale processing - Experience with dbt, Snowflake, and cloud-based ETL/ELT - Familiarity with AWS: S3, MWAA, ECS Fargate, EMR, RDS, Bedrock - Strong data intuition and ability to work independently in ambiguous environments - Experience with test suites — pytest, Great Expectations, dbt tests - Active use of AI coding tools (Cursor, Claude Code, etc.) ### Nice to Have - Healthcare data experience (HIPAA, FHIR, HL7, OMOP) - Data science background (statistical modeling, ML pipelines, feature engineering) - DevOps exposure (CI/CD, Docker, Terraform/CDK) ## Tech Stack - Cloud: AWS (MWAA, ECS Fargate, EMR, RDS, S3, Bedrock) - Data Warehouse: Snowflake - Transformation: dbt - Processing: PySpark, Python - Orchestration: Apache Airflow (MWAA) - AI Tools: Cursor, Claude Code ## Hiring Process - Application Review - Take-Home Assignment (~3 hours, real clinical data) - Technical Interview - Managerial Interview ## Why Century Health - Hard data problems with real clinical impact - Small, high-ownership team — your work ships and matters - Modern cloud-native stack with architectural influence - AI-first engineering culture
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗