Site Reliability Engineer
Core
Design, deploy, and maintain highly available, reliable, and performant production systems and APIs across AWS and GCP, focusing on automation, observability, and disaster recovery.
Role type
Senior Site Reliability Engineer (IC)
Builds
Resilient cloud-native services, automation tools, and disaster recovery solutions
Domain
Cloud Infrastructure / Site Reliability Engineering
Required skills
Python or Go, Kubernetes, Terraform, AWS or GCP, Datadog, SLI/SLO frameworks, FinOps, Disaster Recovery, Service Meshes, HAProxy/NGINX, CI/CD
Preferred skills
GitHub Actions or GitLab Pipelines, Chaos Engineering, HashiCorp Vault, DevSecOps
Technologies
AWS, GCP, Kubernetes, Terraform, Datadog, Argo CD, Kargo, Istio, HAProxy, NGINX, GitHub Actions, GitLab Pipelines, HashiCorp Vault
Responsibilities
Design and maintain highly available production systems and APIs; Define and operationalize SLIs, SLOs, and error budgets; Build end-to-end observability using Datadog; Implement actionable monitoring around Golden Signals; Participate in on-call rotations and incident response; Manage production Kubernetes environments and GitOps workflows; Provision and secure multi-cloud infrastructure with Terraform; Develop automation tools in Python or Go; Monitor cloud infrastructure costs and optimize resources; Configure and troubleshoot service meshes and high-availability proxy solutions. (via careerplan.io/jobs/a9566f08-4aaa-4f65-b50a-83efd309a843-site-reliability-engineer-at-jobgether)
Seniority
Senior, hands-on IC