Senior Site Reliability Engineer (SRE Team)
Core
Own the reliability of Semrush's robust infrastructure and applications by identifying failure points, implementing resilience solutions, and leading engineering practices.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Scalable, reliable, and efficient system architecture; full-stack platform solutions; sophisticated automation tooling in Go/Python.
Domain
SaaS / Brand visibility platform / Cloud infrastructure
Required skills
Kubernetes, Cloud providers, Python, Go, Application failure analysis, Metrics-based debugging, Observability (traces), SLO definition, System architecture design, Incident management, Mentoring
Preferred skills
GCP knowledge
Technologies
Kubernetes, Python, Go, GCP
Responsibilities
Lead changes in common engineering practices; Induce and recover from application failures; Debug applications using metrics and traces; Establish and refine SLOs, cost dashboards, and security hardening initiatives; Collaborate with development teams to design scalable system architecture; Design full-stack platform solutions from concept to production; Build sophisticated tooling in Go/Python to automate operations; Mentor engineers, interview candidates, and lead critical incidents; Manage on-call rotation.
Seniority
Senior, hands-on IC with mentorship responsibilities