Data Engineer
Core
Designing and building real-time data ingestion pipelines and maintaining ETL/ELT infrastructure to support data lakehouses and analytics in the life sciences domain.
Role type
Senior Data Engineer (Cloud/AWS)
Builds
Scalable data pipelines, data lakehouses, and optimized data storage/retrieval systems for enterprise analytics.
Domain
Life Sciences / Cloud Data Engineering
Deliverable
production ML models | infrastructure
Required skills
Python, PySpark, Scala, SQL, AWS services (Glue, Lambda, S3, Redshift, Athena), Databricks, CloudFormation, boto3, GitHub Workflows
Preferred skills
Tableau, Cloudera Data Platform, PyTorch, Pandas, Azure, GCP
Technologies
AWS, Databricks, CloudFormation, GitHub, boto3, Redshift, Athena, Glue, Lambda, S3, Cloudera, Tableau
Responsibilities
Develop and maintain ETL/ELT pipelines for data ingestion; optimize data storage and retrieval for performance and scalability; collaborate with data architects, analysts, and scientists to support data needs; ensure data quality through validation and testing; implement security protocols for sensitive data; design optimal data pipeline architecture; automate manual processes and improve infrastructure scalability.
Seniority
Senior, hands-on IC