Senior Site Reliability Engineer (SRE)
Core
Lead initiatives to maintain and improve the reliability, scalability, and efficiency of production systems through automation and observability.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Automated workflows, deployment safety checks, auto-remediation processes, and observability frameworks.
Domain
Cloud infrastructure and distributed systems
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Cloud platforms (AWS, GCP, Azure), containerization (Docker, Kubernetes), Infrastructure as Code (Terraform, Ansible, Helm), CI/CD pipelines, monitoring tools (Prometheus, Grafana, ELK/EFK, Datadog), distributed systems design, networking, fault-tolerant architecture, scripting (Python, Go, Bash, Java), incident diagnosis, SLO/SLI management, capacity planning, disaster recovery design.
Preferred skills
Technical leadership, mentorship, operational excellence, adaptability to evolving technical landscapes.
Technologies
AWS, GCP, Azure, Docker, Kubernetes, Terraform, Ansible, Helm, Prometheus, Grafana, ELK, EFK, Datadog, Python, Go, Bash, Java
Responsibilities
Lead initiatives to improve system reliability, performance, and scalability; Design and implement automated workflows and auto-remediation processes; Define, enforce, and monitor SLOs/SLIs; Lead incident response, perform postmortems, and implement preventive measures; Mentor junior SREs and promote reliability best practices; Maintain and enhance operational runbooks; Contribute to capacity planning, scaling strategies, and disaster recovery design.
Seniority
Senior, hands-on IC with technical leadership