Principal Data Engineer
Core
Architect and optimize scalable data infrastructure and pipelines to power data-driven decision-making for life-changing medicines.
Role type
Principal or Staff Data Engineer
Builds
Secure, scalable data pipelines (ETL/ELT), data lake and warehouse solutions, and fault-tolerant systems for real-time and batch processing.
Domain
Biotechnology / Life Sciences / Regulated Manufacturing
Deliverable
production ML models | infrastructure
Required skills
Python, Java, Scala, SQL, NoSQL, Hadoop, Spark, Kafka, Flink, Airflow, Luigi, Prefect, containerization, data modeling, schema design, observability, CI/CD
Preferred skills
GenAI, MCP, orchestration platforms, Biotech Enterprise Systems (MES, LIMS, QMS), cloud certifications
Technologies
AWS, Azure, Hadoop, Spark, Kafka, Flink, Airflow, Luigi, Prefect
Responsibilities
Architect and optimize secure, scalable data pipelines; Design data lake and warehouse solutions; Monitor pipeline performance and implement observability; Ensure adherence to data governance and regulatory requirements (21 CFR Part 11, GxP); Define technical standards and design complex data engineering solutions; Mentor junior engineers and serve as a technical leader.
Seniority
Principal (8+ years exp, 1+ biotech/pharma, 3+ cloud) or Staff (10+ years exp, 3+ biotech/pharma, 5+ cloud)