Data Engineer och Systemutvecklare
Core
Building and maintaining a metadata-driven ETL framework and Data Lake platform for clinical research and healthcare decision-making.
Role type
Senior Data Engineer (Healthcare Research)
Builds
Scalable data pipelines, backend services, and integrations for clinical data platforms.
Domain
Healthcare research and clinical data infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Python, ETL/ELT development, event-driven architecture, container platforms (Kubernetes/OpenShift), distributed data processing (Apache Spark), Apache Iceberg, Trino, PII anonymization, SQL databases, object storage
Preferred skills
Experience with public-funded healthcare research data, production monitoring/logging (Datadog, CloudWatch), test/quality assurance frameworks (pytest, Robot Framework, dbt), secrets management, audit logging
Technologies
Apache Iceberg, Spark, Trino, OpenShift, Kafka, RabbitMQ, Python
Responsibilities
Develop and maintain metadata-driven ETL pipelines; Build backend services and integrations with clinical systems; Manage Data Lake architecture; Ensure compliant handling and anonymization of sensitive clinical data; Implement robust solutions on OpenShift with automation and TDD; Collaborate with developers, product owners, and architects.
Seniority
Senior, hands-on IC