CareerPlanGet AI match score →

Site Reliability Engineer - Datadog, Kafka (d/f/m)

Munich💼 Full-time🗓 2026-04-23 → 2026-07-31

Core

Design, build, operate, monitor, and scale infrastructure and event streaming systems (Kafka/CDC) for a SaaS HR platform.

Role type

Senior Site Reliability Engineer (SRE)

Builds

Cloud platform infrastructure, observability stacks, automated runbooks, and event streaming services.

Domain

SaaS HR technology, distributed systems, cloud infrastructure

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Java, Kotlin, distributed systems, Kubernetes, Docker, IaC, Kafka, Datadog, observability, incident management, automation, SLO definition, chaos testing

Preferred skills

CI/CD tooling, JVM tuning, Node.js runtime tuning, AWS MSK Connect

Technologies

Kafka, Datadog, Kubernetes, Docker, Java, Kotlin, Typescript, Python, AWS MSK Connect

Responsibilities

Design and operate event streaming and CDC stacks; create observability dashboards and define SLOs; manage on-call rotations and incident response; automate toil through runbooks and playbooks; mentor peers on reliability practices; conduct chaos testing and root cause analysis.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗