Data Engineer (Public Trust Required)
Core
Design, develop, and maintain scalable, cloud-based data pipelines for large-scale healthcare datasets, specifically cancer registry and population-based data systems.
Role type
Data Engineer (Healthcare/Public Trust)
Builds
Production databases, analytic datasets, and reporting tables for statistical analysis and surveillance.
Domain
Healthcare / Public Health Informatics / Cancer Registry
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Python, PySpark, Spark SQL, pandas, NumPy, PyArrow, matplotlib, Plotly, AWS/Azure/GCP, SAS, R, data quality validation, XML/XSLT
Preferred skills
Migrating SAS-based pipelines to Python/PySpark, cloud-native architectures, serverless functions, containerized applications, Git/SVN
Technologies
Python, PySpark, Spark SQL, pandas, NumPy, PyArrow, matplotlib, Plotly, AWS, Azure, Google Cloud, SAS, R, XML, XSLT
Responsibilities
Design and maintain scalable cloud-based data pipelines; Process and standardize large-scale healthcare datasets; Develop and maintain production databases and reporting tables; Perform rigorous data quality validation and consistency checks; Collaborate with epidemiologists and statisticians to design data solutions; Troubleshoot data pipeline failures and large-scale processing issues
Seniority
Mid-level, hands-on IC