Staff Data Engineer
Core
Building data pipelines, models, and AI infrastructure to ingest, normalize, and govern real-world clinical data from 80+ trial sites to power predictive intelligence for clinical research.
Role type
Staff Data Engineer (Healthcare/AI Infrastructure)
Builds
Data pipelines, data models, and AI infrastructure for clinical research data
Domain
Healthcare technology / Clinical research / AI
Deliverable
production ML models | infrastructure
Required skills
Data engineering, healthcare data (HL7, FHIR, EHR), data modeling from heterogeneous sources, AI/LLM application to data problems, SQL, Python/Java/Scala/Go, distributed processing, streaming, cloud-native storage, orchestration, build vs buy strategy
Preferred skills
ML model training infrastructure, clinical trial operations/EDC systems, HIPAA/SOC 2 compliance, building reliable systems in difficult environments
Technologies
SQL, Python, Java, Scala, Go, distributed processing, streaming, cloud-native storage, orchestration, transformation frameworks
Responsibilities
Own data layer architecture and infrastructure decisions; Build and operate pipelines for data ingestion, normalization, and enrichment; Own data quality and observability systems; Partner with ML teams to define training data requirements and build foundations; Define data governance across the system; Evaluate and select data stack tools; Shape engineering culture and technical decision-making
Seniority
Staff, hands-on IC with strategic influence