Staff Site Reliability Engineer
Core
Own and maintain highly available production systems, lead incident response, and drive operational excellence for scalable cloud infrastructure.
Role type
Staff Site Reliability Engineer
Builds
Scalable cloud infrastructure on AWS, CI/CD and GitOps pipelines, and observability stacks
Domain
Cloud Infrastructure / DevOps
Deliverable
infrastructure
Required skills
AWS, Terraform, Kubernetes (EKS), GitLab CI, ArgoCD, Linux, cloud networking, Python, Bash, Go, observability tools (ELK, Prometheus)
Preferred skills
Engineering degree, cloud certifications (AWS SysOps/DevOps, CKA/CKAD), ITSM/change management experience
Technologies
AWS, Terraform, Kubernetes, EKS, GitLab CI, ArgoCD, PagerDuty, Zenduty, Prometheus, Grafana, ELK, Datadog
Responsibilities
Lead incident response (P1/P2) and conduct RCA/PIRs; Design, build, and manage scalable cloud infrastructure; Develop and optimize CI/CD and GitOps pipelines; Manage observability and on-call operations; Mentor team members and contribute to cloud architecture
Seniority
Staff, hands-on IC with mentorship