Senior Site Reliability Engineer
Core
Build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms for a multi-brand restaurant company.
Role type
Senior Site Reliability Engineer (IC)
Builds
Production-ready, self-healing systems, monitoring solutions, and CI/CD pipelines for high-availability transactional services.
Domain
Cloud-native infrastructure, distributed systems, and high-traffic digital platforms.
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE principles (SLIs, SLOs, Error Budgets), Kubernetes, containerized workloads, incident management, root cause analysis, automation, Infrastructure as Code, load testing, distributed systems architecture, observability strategy, cloud platform expertise (Azure/AWS/GCP), programming (Python/Go/Java/Node.js).
Preferred skills
Chaos engineering, high-volume transactional systems, AI-assisted observability, internal SRE tooling development.
Technologies
Kubernetes, Terraform, Bicep, Python, Go, Java, Node.js, Azure, AWS, GCP.
Responsibilities
Define and manage SLIs/SLOs/Error Budgets; lead high-severity incident response and blameless postmortems; design monitoring, alerting, logging, and tracing solutions; automate manual operational work and build self-healing systems; conduct capacity planning and failure mode analysis; partner with engineering to implement resiliency patterns.
Seniority
Senior, hands-on IC