Senior Site Reliability Engineer (f/m/d)
Core
Design, build, operate, monitor, and scale infrastructure for a SaaS HR platform serving 1.5 million employees across 15,000+ customers.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Cloud platform infrastructure, observability stacks, automated runbooks, and event streaming/CDC systems.
Domain
HR Technology / SaaS / Distributed Systems
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Java, Kotlin, TypeScript, Python, Infrastructure as Code (IaC), Docker, Kubernetes, Kafka, Datadog, System Design, Incident Management, Automation, Observability
Preferred skills
CI/CD tooling (GitHub Actions/GitOps), JVM tuning, Node.js runtime tuning, AWS MSK Connect
Technologies
Java, Kotlin, TypeScript, Python, Docker, Kubernetes, Kafka, Datadog, AWS MSK Connect, GitHub Actions
Responsibilities
Design and improve the full service lifecycle from deployment to continuous improvement; Operate and maintain live services with observability stacks; Ensure sustainable scalability through automation; Collaborate on SLOs, error budgets, and reliability strategies; Support incident management and post-mortems; Reduce toil via process automation and playbooks; Own and maintain the event streaming and CDC stack; Mentor peers on reliability best practices.
Seniority
Senior, hands-on IC