CareerPlanSign in

Data Engineer

London, UK💼 Full-time🗓 2026-05-13 → 2026-09-25

Core

Build pipeline infrastructure and tooling to bridge raw chemical data and machine learning models for materials discovery.

Role type

Senior Data Engineer (Scientific Data Infrastructure)

Builds

Scalable data ingestion pipelines, automated workflows for crystallographic and molecular data, and self-serve data platforms for scientists.

Domain

AI-driven materials science and computational chemistry

Deliverable

production ML models

Required skills

Python, large-scale data processing, workflow orchestration (Airflow/Prefect/Dagster/Flyte), containerization (Docker/Kubernetes), CI/CD, DevOps practices, database management

Preferred skills

MLOps practices, scientific computing data handling, materials science/chemistry domain knowledge

Technologies

Python, Airflow, Prefect, Dagster, Flyte, Docker, Kubernetes

Responsibilities

Design and build robust data pipelines for materials science datasets and computational chemistry outputs; integrate diverse data sources including databases, literature, patents, and lab instruments; implement automated quality checks and monitoring systems for data integrity; collaborate with ML researchers to align data pipelines with model training requirements; support real-time data needs for AI-driven experiments.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.