Site Reliability Engineer (SRE) – II
Core
Lead incident response, automate infrastructure, and drive continuous improvement for critical systems to ensure resilience and scalability.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Automated infrastructure pipelines, monitoring dashboards, and incident response processes
Domain
Financial services / Cloud infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Incident response leadership, Infrastructure as Code (Terraform, Ansible, CloudFormation), Observability (Prometheus, Dynatrace, Splunk), Scripting (PowerShell, Bash, Python), Cloud platforms (AWS, GCP), .NET and Spring Boot support, Hybrid deployment management
Preferred skills
Log analysis (SQL, Splunk), OpenShift and Windows Server experience, Customer-focused mindset
Technologies
Terraform, Ansible, CloudFormation, Prometheus, Dynatrace, Splunk, PowerShell, Bash, Python, AWS, GCP, OpenShift, Windows Server, .NET, Spring Boot
Responsibilities
Lead real-time troubleshooting for high-impact production issues, Build and maintain automation to eliminate manual tasks, Build and optimize monitoring dashboards, Mentor junior SREs and support staff, Drive improvements in deployment and incident response processes, Collaborate across IT and engineering teams to resolve incidents
Seniority
Senior, hands-on IC