Senior Site Reliability Engineer
Core
End-to-end ownership of reliability and operational performance for critical production systems in a high-throughput platform.
Role type
Senior Site Reliability Engineer (IC)
Builds
Secure, resilient, cost-efficient infrastructure and developer-facing tooling
Domain
Cloud infrastructure, distributed systems, observability
Deliverable
production ML models | product features | infrastructure
Required skills
Cloud infrastructure (Kubernetes/EKS, networking, load balancing), Infrastructure as Code (Terraform), Programming (Go, Python), Observability (Datadog, Prometheus, Grafana, OpenTelemetry), Redis/ElastiCache management, Distributed systems failure modes, Incident response, Capacity planning, SLO/SLI/Error budget management, Chaos engineering
Preferred skills
AI tools integration, Progressive delivery practices
Technologies
Terraform, Kubernetes, EKS, Datadog, Prometheus, Grafana, OpenTelemetry, Redis, ElastiCache, Go, Python
Responsibilities
Define and maintain SLIs, SLOs, and error budgets; Lead high-severity incident response and postmortems; Design and operate resilient infrastructure with focus on failure modes; Perform capacity and performance analysis; Build production software and tooling to reduce operational toil; Conduct controlled failure testing and game days; Mentor engineers and partner on production readiness
Seniority
Senior, hands-on IC