Lead Data Engineer
Core
Own the complete lifecycle of enterprise data pipelines from development to production, including roadmap planning, scalable ETL architecture, and secure PHI/PII handling in healthcare.
Role type
Lead Data Engineer
Builds
Scalable ETL pipelines, AI-assisted mapping automation, and RAG-enabled data solutions for healthcare data.
Domain
Healthcare, Regulated Data, Cloud Data Platforms
Deliverable
production ML models
Required skills
PySpark, Python, Advanced SQL, ETL design, AWS data services, PHI/PII security, Data modeling, CI/CD, Data governance, Healthcare data standards
Preferred skills
AI-assisted mapping automation, LLMs for data cleaning, RAG patterns, Vector databases, Terraform, CloudFormation, Databricks, Snowflake
Technologies
AWS (S3, Glue, Lambda, Step Functions, ECS, DynamoDB, Redshift, RDS/PostgreSQL), SQL Server, GitHub, PySpark, Python, SQL
Responsibilities
Design and deploy scalable ETL pipelines using PySpark/Python/SQL on AWS; Own pipeline lifecycle from requirements to production support; Build ingestion pipelines for flat files, APIs, and healthcare sources; Implement AI/LLM-assisted mapping and data quality checks; Ensure secure handling of PHI/PII with encryption and access controls.
Seniority
Lead, hands-on technical design & mentorship