Data Engineer, Clinical Operations
Core
Design, build, and maintain scalable, production-grade data pipelines and GenAI-powered applications for cross-study clinical operations and specimen management workflows.
Role type
Senior Data Engineer (Clinical Operations & GenAI)
Builds
Scalable ETL/ELT pipelines, data models, and GenAI/LLM applications for cross-study and specimen data assets.
Domain
Life Sciences / Clinical Research / Biobanking
Deliverable
production ML models | product features | infrastructure
Required skills
Python, SQL, PySpark, Databricks (Delta Lake, Unity Catalog, Mosaic AI, MLflow), GenAI frameworks, RAG, LLM architectures, cloud-native data platforms, DevOps, semantic modeling
Preferred skills
Databricks certification, life sciences R&D domain knowledge, clinical trial operations experience
Technologies
Databricks, Delta Lake, Unity Catalog, Mosaic AI, MLflow, Python, SQL, Spark, GenAI, RAG, LLMs
Responsibilities
Design and build scalable ETL/ELT pipelines and data models for large, complex cross-study datasets; Develop and operationalize cloud-based GenAI and LLM-powered applications using RAG and fine-tuning; Optimize data platform components for performance, scalability, and cost-effectiveness; Enforce data governance and manage metadata using Unity Catalog; Serve as a technical expert providing guidance on Databricks and technical best practices.
Seniority
Senior, hands-on IC