Senior Site Reliability Engineer
Core
Design, implement, maintain, and oversee reliable production systems at scale, leading incident response and post-incident reviews.
Role type
Senior Site Reliability Engineer (IC)
Builds
Reliable production systems, observability tooling, and vendor integrations at scale
Domain
Cloud infrastructure, DevOps, and enterprise-scale platform engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Cloud platforms (AWS, Azure, GCP), Infrastructure as Code (Terraform, Crossplane), Container orchestration (Kubernetes), Scripting (Python, Ruby, Go), Monitoring (Grafana, New Relic, Datadog), Incident management, SLO governance, Security practices
Preferred skills
Resilient architecture design, Supply chain security, Cross-functional collaboration, Mentoring
Technologies
AWS, Azure, GCP, Datadog, Grafana, Kubernetes, Python, Ruby, Go, Terraform, Crossplane, New Relic
Responsibilities
Design and implement reliable production systems; Lead incident response and post-incident reviews; Proactively track performance and investigate system failures; Create and support observability and monitoring tooling; Foster a reliability-first culture; Coach and mentor other engineers; Collaborate with engineering teams to advance operational excellence
Seniority
Senior, hands-on IC