Software Engineer
Core
Designing, developing, and maintaining scalable data pipelines for ingestion, transformation, and distribution to support reporting, analytics, and AI/ML initiatives.
Role type
Data Engineer
Builds
Scalable ETL/ELT pipelines, data warehouses, data lakehouses, and data models for decision-making and analytics.
Domain
Pharmaceutical industry, data management, and AI/ML support
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Python, SQL, ETL/ELT pipeline design, data warehouse/lakehouse architecture, Apache Spark/PySpark, cloud data platforms (Azure/AWS/GCP), data quality controls, data lineage
Preferred skills
Data governance and compliance (GDPR, GxP), workflow orchestration, DevOps/CI/CD, MLOps, generative AI data patterns (LLM, RAG, embeddings), Docker/Kubernetes
Technologies
Python, SQL, Apache Spark, PySpark, Azure, AWS, Google Cloud Platform, Docker, Kubernetes, Git
Responsibilities
Design and maintain scalable data pipelines for ingestion, transformation, and distribution; Integrate heterogeneous data sources from enterprise systems (ERP, MES, APIs); Optimize data flows for reporting, analytics, and AI/ML use cases; Define and maintain data models, schemas, and metadata; Monitor production pipelines for performance and reliability; Implement automatic data quality controls and lineage tracking
Seniority
Mid-level to Senior, hands-on IC