Senior Data Engineer
Core
Design, build, and maintain scalable data pipelines and infrastructure on AWS to support analytics, machine learning, and Generative AI initiatives for commercial pharmaceutical clients.
Role type
Senior Data Engineer (Pharma/Healthcare Domain)
Builds
End-to-end data pipelines, data processing workflows, and AI/ML data preparation layers on AWS cloud and lakehouse architectures.
Domain
Commercial pharmaceutical data analytics and AI/ML infrastructure
Deliverable
production ML models | infrastructure
Required skills
AWS cloud services (S3, Glue, Lambda, Redshift), Databricks, Apache Spark, SQL, Apache Airflow, data modeling, data lake/lakehouse architecture, commercial pharmaceutical data sources (Xponent, Veeva, MMIT, Plantrak), pharmaceutical commercial data processes (Alignment, Allocation, Split credits, Market basket, Customer universe), pharma KPIs and metrics
Preferred skills
Experience supporting commercial pharmaceutical/healthcare data environments, LLM-based application support, vector embeddings, knowledge retrieval/RAG solutions, legacy system migration
Technologies
Amazon S3, AWS Glue, AWS Lambda, Amazon Redshift, Databricks, Apache Spark, Apache Airflow, Xponent, Veeva, MMIT, Plantrak
Responsibilities
Design and deploy end-to-end data pipelines on AWS; Build and optimize data pipelines for pharma KPIs and reporting; Develop Airflow workflows for orchestration and automation; Integrate and process commercial pharmaceutical data sources; Enable data pipelines for AI/ML and Generative AI workloads; Monitor and troubleshoot pipeline performance and reliability; Support migration to modern cloud and lakehouse architectures
Seniority
Senior, hands-on IC