Staff Site Reliability Engineer (d/f/m)
Core
Design, build, operate, monitor, and scale infrastructure for a SaaS HR platform serving 15,000+ customers and 1.5 million employees.
Role type
Staff Site Reliability Engineer
Builds
Cloud platform infrastructure, shared frameworks, observability stacks, and automation tools for engineering teams.
Domain
HR Technology / SaaS / Distributed Systems
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Java, Kotlin, TypeScript, Python, Infrastructure as Code, Docker, Kubernetes, System Design, Observability, Incident Management, Root Cause Analysis, Process Automation, Chaos Engineering, Mentoring
Preferred skills
CI/CD tooling, JVM tuning, Node.js runtime tuning, Event-driven architectures (Kafka, SNS/SQS)
Technologies
Datadog, AWS, GitHub Actions, GitOps tools, Kafka, SNS, SQS
Responsibilities
Engage in full service lifecycle from design through deployment and operation; Prepare services for production via system design reviews and capacity planning; Operate and monitor live services with observability stacks; Ensure scalability through automation and continuous evolution; Collaborate on defining SLOs and error budgets; Lead incident management including on-call rotations and post-mortems; Identify and reduce toil via automation and playbooks; Define resilience strategies and implement chaos testing; Mentor engineers on reliability best practices.
Seniority
Staff, hands-on IC with mentorship