Senior Site Reliability Engineer
Core
End-to-end ownership of reliability and operational performance for critical production systems in a high-throughput platform.
Role type
Senior Site Reliability Engineer (IC)
Builds
Production infrastructure, observability tooling, and developer-facing software to reduce operational toil.
Domain
Cloud infrastructure, distributed systems, and high-throughput platforms.
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Cloud infrastructure, Kubernetes/EKS, Terraform, Go/Python, Redis/ElastiCache, observability (Datadog/Prometheus/Grafana), incident response, capacity planning, SLO/SLI management, distributed system failure modes, infrastructure-as-code, secure resilient architecture design, chaos engineering, mentoring.
Preferred skills
AI tools for engineering workflows, progressive delivery practices, game days.
Responsibilities
Define and maintain SLIs/SLOs/error budgets, lead high-severity incident response, design secure resilient infrastructure, perform capacity and performance analysis, manage infrastructure via Terraform, build production software and tooling, conduct controlled failure testing, partner on production readiness, participate in on-call rotation, mentor engineers.
Seniority
Senior, hands-on IC