Research Data Engineer
Core
Design and build scalable data systems to power predictive models of cellular behavior for drug discovery.
Role type
Senior Research Data Engineer (Bioinformatics/ML Infrastructure)
Builds
Cloud-native data lake/lakehouse infrastructure, scalable data pipelines, and workflow orchestration for multi-modal scientific data.
Domain
Biotechnology / Drug Discovery / Single-cell multi-omics
Deliverable
production ML models
Required skills
Python, Cloud-native data infrastructure (AWS S3/GCS), Infrastructure-as-Code (Terraform), Data pipeline design, Distributed data processing (Spark, Polars, Dask, DuckDB), Workflow orchestration (Airflow, Dagster, Prefect), Containerization (Docker, Kubernetes), Columnar/scientific data formats (Parquet, Zarr, TileDB, HDF5)
Preferred skills
Biomedical/genomics data experience (BAM, FASTQ, AnnData, OME-Zarr), Regulated/pharma environments, Data governance (FAIR principles), Feature store implementations
Responsibilities
Design and maintain scalable data pipelines ingesting multi-modal scientific data; Optimize data movement, storage layouts, and access patterns for analytical and ML workloads; Stand up and evolve cloud-native data lake/lakehouse infrastructure; Implement data versioning, lineage, and quality monitoring; Partner with data scientists to ensure fast and seamless data workflows; Build and operate workflow orchestration for production pipelines and large-scale batch jobs; Ensure infrastructure meets security, audit, and governance requirements.
Seniority
Senior, hands-on IC