Data Engineer (Bioinformatics)
Core
Build scalable data pipelines to identify prospective clinical trial participants by transforming genomic and health data.
Role type
Senior Data Engineer (Bioinformatics)
Builds
Re-usable, scalable production pipelines for clinical research recruitment
Domain
Healthcare / Genomics / Clinical Research
Deliverable
production ML models | product features
Required skills
Python, Genomic data processing, Bioinformatics tools (PLINK, bcftools), Workflow management (Nextflow, Airflow), Cloud environments (Azure), Containerization (Docker, Kubernetes), Data transformation (Apache Parquet), Version control (Git)
Preferred skills
Genotyping and imputation experience, VCF/BGEN file standards, Spark/Databricks, GA4GH/FAIR data standards
Responsibilities
Design and build robust data pipelines for participant identification, Develop data transformation logic as code, Create prototypes for complex data transformations, Facilitate adoption of data engineering best practices, Provide technical input on upstream data pipeline specifications, Perform ad-hoc data curation and ETL script development
Seniority
Senior, hands-on IC