Data Engineer
Core
Building a real-time data infrastructure platform (TIDE) for capturing and propagating operational data change events using a Change Data Capture (CDC) paradigm and a Lakehouse Medallion architecture.
Role type
Senior Data Engineer (Infrastructure & Streaming)
Builds
Real-time data pipelines, event streaming services, and advanced analytics data layers.
Domain
Enterprise Data Infrastructure / Big Data / Event Streaming
Required skills
Python, Apache Spark (PySpark), Apache Kafka, SQL, Apache Iceberg, Confluent Debezium, Apache NiFi, Git/CI
Preferred skills
Apache NiFi (Groovy scripting), Open Table Formats (Delta/Hudi), Data Lineage/Governance tools, NoSQL (MongoDB), Distributed Storage (Ozone)
Responsibilities
Develop and maintain CDC ingestion pipelines from source databases to the Bronze layer; Process and evolve data flows between Bronze, Silver, and Gold layers using Spark and NiFi; Manage Apache Iceberg tables on Ozone storage including schema evolution; Configure and monitor Kafka topics, offsets, and dead-letter queues; Implement monitoring, alerting, and performance tuning for data pipelines; Enforce project standards for naming, lineage, and access control; Support testing, versioning, and CI/CD activities.
Seniority
Senior, hands-on IC