Site Reliability Engineer
Core
Safeguard the availability, performance, and resilience of client-dependent platforms by owning incident response, sharpening alerting, and growing observability tooling.
Role type
Senior Site Reliability Engineer
Builds
Production platform reliability, incident response processes, and observability tooling
Domain
Cloud Infrastructure / Site Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Incident response and management, Alert tuning and SLI/SLO definition, Observability stack design (metrics, logs, traces), AWS infrastructure operations, Infrastructure as Code (Terraform), Risk identification and remediation
Preferred skills
None stated
Technologies
AWS, Terraform
Responsibilities
Drive response during live incidents and run blameless post-incident reviews, Tune alerting thresholds and codify SLIs/SLOs reflecting customer impact, Build and operate the observability stack for production visibility, Maintain and evolve AWS infrastructure using Terraform, Identify reliability and security risks and follow remediations to production rollout
Seniority
Senior, hands-on IC