Data Engineer
Core
Build and optimize scalable data pipelines and analytical datasets on an AWS-based lakehouse platform to support analytics and operational use cases.
Role type
Data Engineer
Builds
Production-grade data solutions, batch data pipelines, and reusable ingestion/transformation workflows
Domain
Healthcare technology / Data Engineering
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Python, SQL, PySpark/Spark, AWS S3, AWS Glue, AWS Athena, dbt, Airflow, Apache Iceberg, schema evolution, partition optimization, data validation, performance tuning
Preferred skills
Hudi, Delta
Technologies
Apache Iceberg, AWS Glue, Athena, dbt, Airflow, Python, SQL, PySpark, AWS S3
Responsibilities
Develop and maintain batch data pipelines; Build reusable ingestion and transformation workflows; Build and maintain Apache Iceberg datasets; Support schema evolution and partition optimization; Optimize Athena queries and troubleshoot performance bottlenecks; Implement data validation checks and monitoring; Support production issue troubleshooting
Seniority
Mid-level (4–5 years experience)