AI Data Engineer
Core
Design, build, and maintain scalable data pipelines to ingest, process, and store massive and diverse healthcare datasets for LLM and machine learning model fine-tuning.
Role type
Data Engineer (LLM training data lifecycle)
Builds
Scalable data pipelines for healthcare datasets
Domain
Healthcare + Big Data Engineering
Deliverable
production ML models
Required skills
Python, Scala, Java, Apache Spark, ETL, data warehousing, data modeling, cloud data services (AWS/GCP/Azure), data validation, schema design
Preferred skills
FHIR, HL7, data orchestration (Apache Airflow), machine learning concepts
Technologies
Apache Spark, Apache Airflow, AWS, GCP, Azure
Responsibilities
Design and maintain scalable data pipelines for healthcare datasets; implement data validation and monitoring; develop and optimize data structures for LLMs; acquire new data sources ensuring HIPAA compliance; troubleshoot and optimize pipeline performance; document data models and processes.
Seniority
Mid-Senior, hands-on IC