Associcate Data Engineer
Core
Design, build, and maintain scalable data pipelines and ETL/ELT processes to ingest, transform, and process large-scale structured and unstructured data for actionable business insights.
Role type
Associate Data Engineer
Builds
End-to-end data pipelines, data models, and automated workflows on cloud platforms
Domain
Biotechnology / Big Data Engineering
Deliverable
production ML models | product features | dashboards & analysis
Required skills
PySpark, SQL, Databricks, ETL/ELT, data modeling, data governance, cloud platforms (AWS), workflow orchestration, performance tuning
Preferred skills
Python, Apache Airflow, SageMaker, data protection regulations (GDPR/CCPA), multi-source integration (APIs, cloud storage)
Technologies
Databricks, PySpark, SparkSQL, AWS, Apache Airflow, SQL databases, APIs
Responsibilities
Design and develop data pipelines leveraging Databricks and PySpark to ingest and process large datasets; Implement automated workflows for data ingestion and deployment; Optimize Spark job performance through tuning, caching, and partitioning; Collaborate with Data Architects and Data Scientists to design end-to-end solutions; Ensure data quality and security across systems; Manage data pipeline projects from inception to deployment.
Seniority
Mid-level, hands-on IC