Expert Site Reliability Engineer
Core
Drive reliability, availability, scalability, and operational resilience of critical technology services through advanced software engineering, automation, and observability practices.
Role type
Senior IC Site Reliability Engineer
Builds
Critical technology services with high availability and resilience
Domain
Technology services / Cloud infrastructure
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering, DevOps, Platform Engineering, Cloud platforms, Kubernetes, Monitoring, Observability, Alerting, SLIs, SLOs, Automation, Scripting, CI/CD, Infrastructure as Code, Terraform, Incident Management, Root Cause Analysis, Performance Engineering, Capacity Planning
Preferred skills
Banking, FinTech, or Payment environments experience
Technologies
Kubernetes, Terraform
Responsibilities
Define and implement advanced reliability engineering practices; Establish and monitor SLIs, SLOs, and reliability targets; Design automation to reduce manual operational activities; Develop and enhance monitoring, observability, alerting, and incident detection capabilities; Lead technical analysis and resolution of complex production incidents; Conduct root-cause analysis and drive permanent corrective and preventive actions; Design solutions to improve system availability, scalability, capacity, and disaster resilience; Identify reliability risks and recommend architectural and engineering improvements; Drive performance engineering and capacity planning for critical services; Provide advanced technical guidance and mentorship on SRE practices.
Seniority
Senior, hands-on IC with mentorship responsibilities