CareerPlanSign in

Site Reliability Engineer

🌐 Remote💼 Full-time🗓 2026-06-24

Core

Safeguard the availability, performance, and resilience of client-dependent platforms by owning incident response, sharpening alerting, and growing observability tooling.

Role type

Senior Site Reliability Engineer

Builds

Production platform reliability, incident response processes, and observability tooling

Domain

Cloud Infrastructure / Site Reliability Engineering

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Incident response and management, Alert tuning and SLI/SLO definition, Observability stack design (metrics, logs, traces), AWS infrastructure operations, Infrastructure as Code (Terraform), Risk identification and remediation

Preferred skills

None stated

Technologies

AWS, Terraform

Responsibilities

Drive response during live incidents and run blameless post-incident reviews, Tune alerting thresholds and codify SLIs/SLOs reflecting customer impact, Build and operate the observability stack for production visibility, Maintain and evolve AWS infrastructure using Terraform, Identify reliability and security risks and follow remediations to production rollout

Seniority

Senior, hands-on IC

Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.