Site Reliability Engineering (SRE ) lead
Core
Lead troubleshooting, root cause analysis, and automation for complex production incidents and cloud platform reliability.
Role type
Senior SRE Lead (IC with team leadership)
Builds
Production-ready cloud infrastructure, automated operational runbooks, and self-healing capabilities
Domain
Financial Services / Cloud Infrastructure & Distributed Systems
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Incident management, Root cause analysis (RCA), Infrastructure as Code (IaC), CI/CD pipeline design, Cloud platform operations, Team leadership, Operational metrics analysis
Preferred skills
Distributed systems architecture, Automation framework development, Stakeholder management, Change management methodologies
Technologies
AWS, Azure, Kubernetes, Docker, Python, PowerShell, Shell, GitHub Actions, Azure DevOps, Jenkins, GitLab, Datadog, Splunk, Dynatrace, Grafana, Prometheus, CloudWatch, Azure Monitor, OpenTelemetry, ServiceNow, Jira, Terraform, Ansible, SQL
Responsibilities
Lead troubleshooting and resolution of complex production incidents; Conduct comprehensive root cause analysis and implement corrective actions; Design and enhance monitoring, observability, and alerting systems; Drive automation initiatives using scripting and IaC; Partner with engineering teams to remediate recurring reliability issues; Serve as Incident Commander during major outages; Provide leadership, coaching, and mentoring for SRE and production support engineers
Seniority
Senior, hands-on IC with team leadership