Sr Data Engineer
Core
Design, develop, and optimize complex data pipelines, integration frameworks, and metadata-driven architectures to enable seamless data access, self-service analytics, and AI-driven insights for R&D in biotech/pharma.
Role type
Senior IC data engineer (big data & distributed computing)
Builds
Scalable ETL/ELT pipelines, real-time and batch data processing solutions, governed data fabric architectures, and CI/CD pipelines for automated deployments.
Domain
Biotechnology / Pharmaceuticals / Big Data Engineering
Deliverable
production ML models | product features | infrastructure
Required skills
Databricks, PySpark, SparkSQL, AWS, Python, SQL, workflow orchestration, performance tuning, data modeling, data governance, CI/CD, metadata management, data lineage tracking, data virtualization, RBAC implementation, distributed computing frameworks (Apache Spark, Hadoop)
Preferred skills
Biotech & Pharma domain expertise, API development, vector databases for LLMs, OLAP/OLTP database optimization, Git, Jenkins, Maven, automated unit testing
Technologies
Databricks, PySpark, SparkSQL, AWS, Apache Spark, Hadoop, Jenkins, Maven, Git, SQL, NOSQL, vector databases
Responsibilities
Design and maintain scalable ETL/ELT pipelines for structured, semi-structured, and unstructured data; Implement real-time and batch data processing solutions integrating multiple sources; Optimize big data processing frameworks for high availability and cost efficiency; Work with metadata management and data lineage tracking tools; Ensure data security, compliance, and role-based access control; Optimize query performance, indexing, partitioning, and caching; Develop CI/CD pipelines for automated deployments; Implement data virtualization techniques; Collaborate with data architects, business analysts, and DevOps teams.
Seniority
Senior, hands-on IC