Senior Data Engineer — Translational Data Products
Core
Build, operate, and own end-to-end data products (biomarker, biospecimen, clinical trial, omics) that make R&D data usable and AI-ready for translational scientists and analysts.
Role type
Senior Data Engineer (Translational Data Products)
Builds
Governed, curated, linked, documented, AI-ready data products and datasets
Domain
Life sciences / Biopharmaceutical R&D
Deliverable
production ML models | product features
Required skills
SQL (complex joins, window functions, performance tuning), Python, AWS (S3, IAM, Glue, Lambda, ECS, Fargate), Docker, Infrastructure as Code (AWS CDK, CloudFormation), GitHub Actions, Test-driven development, Life sciences data (clinical, molecular, assay), Stakeholder engagement with scientists
Preferred skills
Translational medicine/biomarker domain knowledge, Clinical trial data standards (SDTM/ADaM), Databricks (Unity Catalog, Delta Lake), Omics modality depth, Claude Code extension experience
Technologies
Databricks, AWS, Unity Catalog, GitHub Actions, AWS CDK, CloudFormation, Docker, NVIDIA clusters, Wiz, Dependabot
Responsibilities
Design and maintain ETL/data normalization pipelines on Databricks and AWS; Unlock data access by working with data owners to bring new sources under governance; Engage stakeholders to translate scientific questions into robust data product designs; Translate between business/science and data by creating data models and SQL queries; Validate data products with test-driven and validation-driven approaches; Engineer for robustness using infrastructure as code and automated security checks; Work with ML engineers to prepare feature-ready datasets for GenAI workloads
Seniority
Senior, hands-on IC