Principal Software Engineer - Data Platform (Iceberg/Trino)
Core
Design and own the architecture of an on-premise lakehouse engine using Apache Iceberg, Trino, and Spark to power healthcare data platforms.
Role type
Principal Software Engineer (Data Platform Architecture)
Builds
On-premise lakehouse engine, catalog services, and data transformation pipelines for healthcare clients.
Domain
Healthcare technology, distributed systems, open-source data platforms.
Deliverable
production ML models | product features | infrastructure
Required skills
Distributed SQL engines (Trino/Presto, Spark SQL), Apache Iceberg internals, Java/Python development, on-premise data platform design, query planning, performance engineering.
Preferred skills
Iceberg catalog services (Polaris, Nessie), cloud warehouse internals (Snowflake, BigQuery), regulated/air-gapped environment experience, healthcare data experience.
Technologies
Apache Iceberg, Trino, Spark, Java, Python, S3, Kubernetes, VMs.
Responsibilities
Own lakehouse reference architecture (Iceberg, Trino, Spark, storage layout); design on-premise replacements for cloud warehouse capabilities; validate catalog and query engine performance; set platform-wide standards for table layout and maintenance; lead SQL dialect strategy for workload porting; mentor senior engineers; partner on storage sizing and capacity planning.
Seniority
Principal, hands-on IC with mentorship