About the job
Core
Build the biomedical data foundation powering AI-driven drug discovery and translational research by integrating multimodal biomedical data into standardized, AI-ready datasets.
Role type
Senior Biomedical Data Engineer
Builds
Scalable biomedical data pipelines, ontologies, knowledge graphs, and AI-ready datasets for foundation models and digital twins.
Domain
Biomedical data engineering, AI-driven drug discovery, translational research
Deliverable
production ML models
Required skills
Python, SQL, ETL pipeline development, cloud-based data platforms (AWS), data modeling, metadata management, data quality validation, handling structure and unstructured multimodal biomedical data
Preferred skills
Genomics, transcriptomics, proteomics, imaging, biomedical ontologies, knowledge graphs, LLM-based information extraction, foundation model development, AI-driven drug discovery
Technologies
AWS, Python, SQL
Responsibilities
Design and develop biomedical data pipelines integrating internal and external datasets; Develop standardized, AI-ready multimodal biomedical datasets; Design and implement biomedical data models for structured and unstructured data; Develop and optimize ETL workflows; Design biomedical ontologies, metadata standards, and knowledge graphs; Perform data harmonization, normalization, and quality validation across heterogeneous sources; Collaborate with AI scientists to prepare datasets for machine learning and foundation models; Evaluate external biomedical databases and support data acquisition strategies.
Seniority
Senior, hands-on IC