Data Engineer (Forward Deployed)
Core
Build compute, data, and tooling foundations for scientific discovery in health, food security, climate, and AI/robotics, acting as an interface between core systems and research projects. (via careerplan.io/jobs/7c665acd-24c5-42b1-915e-cd87879bb5f5-data-engineer-forward-deployed-at-ellison-institute-of-technology)
Role type
Forward Deployed Data Engineer
Builds
Cloud/GPU infrastructure, data products, shared data models, and reproducible data pipelines for ML and research.
Domain
Life Sciences / Bioinformatics / Scientific Computing
Deliverable
production ML models
Required skills
Python, data storage and manipulation (relational, sharding, indexing), cloud compute platforms, Linux, distributed/parallelized data systems, data engineering best practices (versioning, lineage, schema design), containerization (Docker), orchestration (Kubernetes, Slurm), distributed frameworks (Spark, Ray), bioinformatics tools (NextFlow), ML data preparation.
Preferred skills
Genomics, proteomics, scientific background.
Technologies
Python, Docker, Kubernetes, Slurm, Spark, Ray, LMDB, Arrow, HDF5, fastq, fasta, cif, NextFlow, Oxford, London
Responsibilities
Partner with scientists to deliver robust, reproducible data pipelines; own ingestion, storage, curation, and transformation of diverse biological datasets; package and deploy code in research environments; scale processing across distributed cloud warehouses and GPU compute.
Seniority
Mid-Senior, hands-on IC
