Senior Site Reliability Engineer
Core
Senior SRE owning complex reliability and platform problems, building AI-native workflows, and shaping architecture for a global HR platform.
Role type
Senior IC Site Reliability Engineer (Platform & AI)
Builds
Global HR platform infrastructure, reliability tooling, and AI-assisted operational workflows
Domain
SaaS / HR Tech / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, AWS, Terraform, SLO/SLI/Error Budgets, OpenTelemetry, Grafana/Prometheus, CI/CD, Golang, Bash, AI agentic workflows
Preferred skills
Linux systems, security (defensive/offensive), back-end programming (Elixir/Nodejs/Python)
Technologies
Kubernetes, Docker, AWS, Terraform, OpenTelemetry, Grafana, Prometheus, GitLab CI, GitHub Actions, Golang, Bash
Responsibilities
Lead solution discovery and delivery for reliability and infrastructure problems; Contribute to platform architecture, tooling, and roadmap; Define and operate reliability practices (SLOs/SLIs, alerting, observability); Resolve cross-team requests and turn recurring issues into reusable fixes; Operationalize AI for the team via agentic workflows and tooling; Mentor engineers and participate in hiring/onboarding; Collaborate with Security on platform hardening; Participate in incident response and on-call rotations
Seniority
Senior, hands-on IC with mentorship responsibilities