Site Reliability Engineer Engineer
Core
Build reliable, observable systems for client teams to operate at scale, reducing toil and preventing incidents.
Role type
Senior Site Reliability Engineer (IC)
Builds
Backend services on AWS (ALB, ECS/Fargate, Aurora) and Python services
Domain
Cloud Infrastructure / SRE / Python Backend
Required skills
AWS (ALB, ECS/Fargate, Aurora, Lambda), Python backend engineering, Kubernetes, Docker, Terraform, CI/CD, observability (metrics, logging, tracing), incident response, SLOs/error budgets, Linux/Unix, scripting (Python, Bash, Go)
Preferred skills
Internal developer platforms, chaos engineering, cloud security/compliance, AI infrastructure/LLM serving, FinOps, mentoring
Technologies
AWS, Python, Kubernetes, Docker, Terraform, CloudFormation, Pulumi, Git, Bash, Go
Responsibilities
Design observability systems and define SLOs; own 24/7 P0 on-call rotation; establish reliability standards and SLAs; mentor and onboard SRE engineers; participate in incident response and postmortems; implement reliability improvements and capacity planning; build self-healing automation; support AI application operations; create runbooks and documentation.
Seniority
Senior, hands-on IC with mentorship responsibilities