Associate Data Engineer
Core
Build, maintain, and optimize scalable data pipelines and ETL/ELT workflows to support analytics, reporting, and machine learning use cases.
Role type
Associate Data Engineer
Builds
Reliable datasets and data pipelines for analytics, reporting, and ML
Domain
Biotechnology / Data Engineering
Deliverable
production ML models | product features
Required skills
Python, PySpark, Databricks, SQL, Delta Lake, ETL/ELT, data pipeline development, data quality validation, feature engineering, cloud platforms (AWS/Azure/GCP), Git
Preferred skills
Workflow orchestration, CI/CD, MLflow, scikit-learn, Tableau/Power BI, data governance
Technologies
Databricks, PySpark, Python, Delta Lake, AWS, Azure, GCP, Git, MLflow, scikit-learn, Tableau, Power BI
Responsibilities
Develop, test, and maintain data pipelines; Ingest and transform structured/semi-structured data; Optimize Spark jobs and notebooks; Troubleshoot pipeline failures and performance bottlenecks; Create and maintain documentation for pipelines and data definitions; Monitor scheduled jobs for timely delivery; Support basic AI/ML data preparation and feature engineering
Seniority
Mid-level, hands-on IC