Process Data Engineer I - Pharmaceutical Product Development
Core
Build scientific data products and scalable data pipelines to enable AI, machine learning, and scientific decision-making for pharmaceutical product development.
Role type
Early-career Data Engineer (Pharmaceutical Product Development)
Builds
Data products for molecular features, material properties, lab data, process/manufacturing parameters, and stability/performance data; scalable pipelines for analytics and modeling.
Domain
Pharmaceutical Product Development (Chemical Process, Biologics, Drug Product, Analytical Development)
Deliverable
production ML models | product features
Required skills
SQL, Python, ETL/ELT, Delta Lake, Lakehouse Architecture, Medallion Architecture, Data Modelling, Distributed Computing, Data Validation, Databricks, dbt, AWS, Data Product Design, Vector Databases, Data Quality Engineering, Data Observability, Metadata Management, Master Data Management, Data Lineage, Data Governance, CI/CD, API Integration, Workflow Automation
Preferred skills
PySpark, experience with pharmaceutical product development datasets, familiarity with AI/ML workflows
Technologies
Databricks, dbt, AWS, Delta Lake, Vector Databases, Git
Responsibilities
Build scientific data products; develop and maintain scalable data pipelines; transform raw data into model-ready products; work with cross-functional teams to improve data quality; automate data workflows; contribute to reusable data assets.
Seniority
Early-career, hands-on IC