Staff Site Reliability Expert
Core
Design and maintain scalable, automated AWS infrastructure; build tools and platform services enabling product engineers to ship and operate independently.
Role type
Staff Site Reliability Engineer (IC)
Builds
Scalable AWS infrastructure, platform services, and automation tools
Domain
Cloud Infrastructure (AWS), SaaS
Deliverable
production ML models | infrastructure
Required skills
AWS, Terraform, Docker, Kubernetes, ECS, Linux, Python, Ruby, Go, Shell scripting, observability, incident management, disaster recovery, security practices, cloud cost optimization, datastores (MySQL, PostgreSQL, Redis, DynamoDB)
Preferred skills
Agile, continuous delivery, testing
Technologies
AWS, Terraform, Docker, Kubernetes, ECS, MySQL, PostgreSQL, Redis, DynamoDB
Responsibilities
Design and maintain scalable, automated AWS infrastructure; build tools and platform services; partner with engineering teams to design resilient systems; champion observability and incident management; lead incident response and mentor engineers; participate in on-call rotation
Seniority
Staff, hands-on IC with mentorship