Data Engineer II
Core
Designing, building, and maintaining robust data pipelines and infrastructure for collecting, storing, and processing large datasets to support business and research needs.
Role type
Senior IC data engineer (healthcare domain)
Builds
Scalable data pipelines, unified data repositories (warehouses/lakes), and analytics tools for data scientists and analysts.
Domain
Healthcare / Data Engineering
Deliverable
production ML models | infrastructure
Required skills
Python (PySpark), SQL, Apache Spark, Apache Airflow, Git, Cloud data platforms (Azure/AWS/GCP), Data warehousing
Preferred skills
Healthcare data experience (EHR, REDCap), Cloud certifications (Azure/GCP/AWS)
Technologies
Apache Spark, Apache Airflow, Microsoft Fabric, Azure Data Factory, AWS Glue, Azure Synapse, Google BigQuery, AWS Redshift, PySpark
Responsibilities
Design and optimize ETL processes for structured and unstructured data; Integrate data from disparate systems into unified repositories; Build custom queries, scripts, and dashboards for insight generation; Automate manual workflows and improve data delivery pipelines; Ensure data governance, accuracy, and compliance (e.g., HIPAA); Collaborate with stakeholders to translate requirements into data solutions.
Seniority
Mid-Senior, hands-on IC