CareerPlanGet AI match score →

Site Reliability Engineer

🌐 Remote💼 Full-time🗓 2026-06-24

Core

Safeguard the availability, performance, and resilience of client-dependent platforms by owning incident response, sharpening alerting, and growing observability tooling.

Role type

Senior Site Reliability Engineer

Builds

Production platform reliability, incident response processes, and observability tooling

Domain

Cloud Infrastructure / Site Reliability Engineering

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Incident response and management, Alert tuning and SLI/SLO definition, Observability stack design (metrics, logs, traces), AWS infrastructure operations, Infrastructure as Code (Terraform), Risk identification and remediation

Preferred skills

None stated

Technologies

AWS, Terraform

Responsibilities

Drive response during live incidents and run blameless post-incident reviews, Tune alerting thresholds and codify SLIs/SLOs reflecting customer impact, Build and operate the observability stack for production visibility, Maintain and evolve AWS infrastructure using Terraform, Identify reliability and security risks and follow remediations to production rollout

Seniority

Senior, hands-on IC

Rewrite
## About the role Sigma360 is seeking a Site Reliability Engineer to help safeguard the availability, performance, and resilience of the platform our clients depend on. This is a hands-on engineering role embedded between our DevOps and Security teams: you will own incident response alongside the engineers who build the systems, sharpen our alerting so the right people are paged for the right reasons, and grow the observability and monitoring tooling that lets us understand production behavior at a glance. We are looking for an engineer who treats reliability as a product — proactively hunting down sources of toil and risk before they turn into outages — rather than someone who waits for tickets. This role is fully remote within Europe. You will work closely with engineers across multiple time zones and participate in a healthy, well-staffed on-call rotation that values learning from incidents over blame. ## Key Responsibilities - Partner with our DevOps and Security teams across the full incident lifecycle: drive response during live incidents, run blameless post-incident reviews, and turn each operational surprise into durable improvements in code, runbooks, or tooling. - Improve the signal-to-noise ratio of our alerting by tuning thresholds, removing duplicate or low-value alerts, and codifying SLIs/SLOs that reflect actual customer impact. - Build and operate the observability stack — metrics, logs, traces, and dashboards — so engineers across the company can answer their own questions about production. - Maintain and evolve our AWS infrastructure through Terraform, with a focus on safe, reviewable changes and reproducible environments. - Identify reliability and security risks proactively, propose remediations, and follow them through to production rollout. ## Qualifications - Strong production operations experience. You have supported business-critical systems and can stay calm and methodical during a live incident, then translate what you learned into structural fixes. - Cloud and infrastructure-as-code fluency. You can design, review, and refactor Terraform that manages real AWS workloads, and you understand the trade-offs of the AWS primitives you reach for. - Observability sensibility. You know what makes an alert actionable, what makes a dashboard useful, and how to instrument a service so the next on-call engineer can debug it without paging the author. - Collaborative, low-ego communication. You work well with engineers who are not SREs, write clearly in async channels, and can push back on a design without making it personal. ## Must Have - 4+ years of experience in an SRE, Production Engineering, DevOps, or similar role supporting production systems at meaningful scale - Hands-on experience operating workloads on AWS in production - Practical experience with Terraform, including reviewing and authoring modules that change real infrastructure - Experience participating in a 24/7 on-call rotation, including incident command and post-incident review - Track record of im
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗