Bioinformatics Engineer, London
Core
Building a high-fidelity biological data layer and operating large-scale pipelines to harmonize disparate public and internal datasets for machine learning in drug discovery.
Role type
Senior IC bioinformatics engineer (data platform)
Builds
Production-grade bioinformatics pipelines and curated biological datasets for ML training
Domain
Biopharmaceuticals / Computational Biology / Drug Discovery
Deliverable
production ML models
Required skills
Large-scale bioinformatics data processing (FASTQ, BAM, mzXML), Python software development, automated pipeline creation, database harmonization and versioning, user enablement for research teams
Preferred skills
Domain-specific workflow systems (Nextflow), data orchestration frameworks (Dagster, Apache Beam), GCP infrastructure, high-performance DataFrame libraries (Polars), SQL, machine learning concepts
Technologies
Python, FASTQ, BAM, mzXML, Ensembl, UniProt, Reactome, Open Targets, Nextflow, Dagster, Apache Beam, Google Cloud Platform, Polars, SQL
Responsibilities
Develop and operate large-scale bioinformatics pipelines for high-throughput data analysis; Apply bioinformatics best practices to ingestion and harmonization of complex datasets; Harmonize disparate public databases with rigorous versioning and mapping strategies; Act as a strategic partner to ML Research, Computational Biology, Drug Development, and Chemistry teams; Participate in research projects as a Deployed Engineer providing customised solutions; Provide documentation, guidance, and training on data resources and curation processes
Seniority
Senior, hands-on IC