Data Infrastructure Engineer
Core
Design, build, and maintain data systems connecting nanopore sequencing instruments to analysis, transforming raw instrument output into clean, queryable datasets for life science discovery.
Role type
Senior Data Infrastructure Engineer (Bioinformatics)
Builds
End-to-end data pipelines, centralized data lakes, and self-serve visualization tools for proteome sequencing data.
Domain
Life Sciences / Bioinformatics / Nanopore Sequencing
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Python, SQL, AWS cloud services, Docker, Nextflow/Snakemake, data modeling, ETL/ELT, data visualization tools, machine learning model lifecycle management, Linux/shell scripting
Preferred skills
Nanopore data formats (POD5, FAST5), Seqera Platform, real-time data processing, AI coding assistants, early-stage biotech infrastructure experience
Technologies
AWS, Docker, Nextflow, Seqera Platform, PostgreSQL, DuckDB, BigQuery, Snowflake, Metabase, Dash, Streamer, Looker, Python, Linux
Responsibilities
Own and extend Nextflow pipelines for nanopore sequencing output processing; Design and implement data models and schemas for sequencing data; Build ETL workflows for centralized data lakes; Deploy and maintain data visualization tools for scientists; Manage machine learning classifier model lifecycle; Automate generation of standard analysis outputs and reporting.
Seniority
Senior, hands-on IC