4467- Software Development Engineer-III (Data Engineer) (Iceberg/Trino)
Core
Build and operate data pipelines for an on-premise healthcare platform, ingesting raw data into Apache Iceberg and creating layered tables via Trino SQL for analytics.
Role type
Senior Software Engineer (Data Engineer)
Builds
Production data pipelines, Iceberg tables, and Trino-based analytics layers for healthcare clients.
Domain
Healthcare / Data Engineering
Deliverable
production ML models | product features
Required skills
Apache Spark, Trino/Presto, Apache Iceberg, SQL, Python/Java, S3-compatible storage, workflow orchestration (Airflow), CI/CD
Preferred skills
Healthcare data formats (HL7, CCDA), regulated-environment experience
Technologies
Apache Spark, Trino, Apache Iceberg, Parquet, Airflow, AWS HealthLake, Amazon Bedrock
Responsibilities
Build Spark ingestion jobs for high-volume raw files with schema handling and bad-record quarantine; Develop Trino SQL transform pipelines for validation, business rules, deduplication, and aggregates; Port warehouse SQL workloads to Trino/Spark dialects; Automate Iceberg table maintenance (compaction, snapshot expiry, orphan-file cleanup); Tune query and pipeline performance (partitioning, file sizing, statistics); Instrument pipelines with data-quality checks and alerting.