Site Reliability Engineer - Datadog, Kafka (d/f/m)
Core
Design, build, operate, monitor, and scale infrastructure and event streaming systems (Kafka/CDC) for a SaaS HR platform.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Cloud platform infrastructure, observability stacks, automated runbooks, and event streaming services.
Domain
SaaS HR technology, distributed systems, cloud infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Java, Kotlin, distributed systems, Kubernetes, Docker, IaC, Kafka, Datadog, observability, incident management, automation, SLO definition, chaos testing
Preferred skills
CI/CD tooling, JVM tuning, Node.js runtime tuning, AWS MSK Connect
Technologies
Kafka, Datadog, Kubernetes, Docker, Java, Kotlin, Typescript, Python, AWS MSK Connect
Responsibilities
Design and operate event streaming and CDC stacks; create observability dashboards and define SLOs; manage on-call rotations and incident response; automate toil through runbooks and playbooks; mentor peers on reliability practices; conduct chaos testing and root cause analysis.
Seniority
Senior, hands-on IC