Site Reliability Engineering Lead
Core
Define and implement SRE best practices, establish reliability goals, and lead incident response for mission-critical, high-volume transaction web-based software systems.
Role type
Senior IC SRE Lead (Strategy & Mentorship)
Builds
Self-healing systems, runbook automation, and resilient cloud infrastructure for finance and insurance clients.
Domain
Cloud Operations / Financial Services / Insurance
Deliverable
production ML models | infrastructure
Required skills
SRE best practices, disaster recovery, automation, AWS services, observability tools, incident management, system resilience design
Preferred skills
AIOps, ML-based anomaly detection, geographically distributed team leadership, regulated environment experience
Technologies
AWS, Prometheus, Grafana, ELK, AWS CloudWatch, Gremlin, Chaos Monkey, AWS FIS
Responsibilities
Define and implement SRE best practices across the organization, Establish SLIs, SLOs, and SLAs for critical services, Act as primary escalation point for critical production issues and lead major incident response, Champion monitoring, logging, and alerting strategies, Collaborate with development teams to integrate reliability into application design, Mentor and guide a team of SREs
Seniority
Senior, hands-on IC with leadership responsibilities