Data Engineer (Python)
Required skills
3+ years data engineering experience (level dependent) with real pipeline delivery beyond ad-hoc scripts. Strong Python + SQL; comfortable building transformations, validation tooling, and pipeline glue code. Practical streaming/CDC fundamentals (ordering, duplication, replay, idempotency) and Kafka ecosystem experience. Familiar with lakehouse/storage and query layers (e.g., Hudi/Iceberg/Delta, Trino/Hive/Postgres) and how to make datasets usable. Comfortable working in Kubernetes/container environments and documenting decisions clearly.
Preferred skills
Great Expectations or similar data quality tooling; metadata/lineage platforms (OpenMetadata/DataHub/Atlas). Experience shipping in on-prem or air-gapped environments; governance/policy awareness for regulated customers. German language (B1+) and/or experience with OSINT/GEOINT/multi-INT data shapes.
Technologies
NiFi, Kafka, Kafka Connect/Streams, CDC, Hudi/Iceberg/Delta, Trino/Hive/Postgres, Kubernetes, Python, SQL.
Responsibilities
Prototype ingestion and connector patterns (batch + streaming) using NiFi, Kafka, Kafka Connect/Streams, and CDC approaches. Design "prototype-grade but adoptable" schemas and data models with clear semantics and evolution discipline. Build incremental lakehouse datasets (Hudi/Iceberg/Delta patterns) and produce queryable outputs for realistic latency/throughput evaluation. Bake in data quality and provenance mindset early (checks, metadata hooks, operability basics). Containerize and deploy prototypes on Kubernetes; deliver minimal runbooks/configs that make adoption straightforward. Produce adoption artifacts: schemas, reference implementations, technical design notes, and an integration backlog.
Seniority
level dependent.
Domain
data intelligence, search, ML enrichment, investigative workflows.