Senior Site Reliability Engineer II
Core
Design, build, and operate highly available, scalable systems in AWS, focusing on infrastructure reliability, observability, and incident response.
Role type
Senior Site Reliability Engineer (hands-on IC)
Builds
Production infrastructure, CI/CD pipelines, and operational workflows
Domain
Cloud Infrastructure (AWS), DevOps, Site Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
AWS, Terraform, Linux systems, networking, troubleshooting, incident response, Git workflows, CI/CD pipelines, ITSM workflows, documentation
Preferred skills
SLOs and error budgets, Infrastructure-as-Code best practices, containers and orchestration, large-scale high-availability systems, technical mentoring
Technologies
AWS, Terraform, Grafana, Pingdom, Uptrends, Azure DevOps, GitHub, ServiceNow, Confluence, Jira, Docker, Kubernetes
Responsibilities
Design and operate scalable systems; write and review Terraform; manage monitoring and observability; lead incident response and root cause analysis; define and manage SLOs/SLIs; build CI/CD pipelines; automate operational tasks; mentor engineers
Seniority
Senior, hands-on IC