Lead Site Reliability Engineer, Public Cloud
Core
Lead SRE driving cross-team standardization, reliability improvements, and AI/LLM integration for public cloud platforms.
Role type
Senior IC Lead Site Reliability Engineer (Public Cloud)
Builds
Production-grade reliability standards, automation pipelines, and AI-enhanced operational workflows for AWS, Azure, and GCP.
Domain
Financial Services / Public Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
SRE, SLOs/SLIs, error budgets, observability, incident response, IaC, AWS, Azure, Python, Go, Bash, Terraform, CI/CD, AIOps, LLMs
Preferred skills
AI-assisted troubleshooting, agentic workflows, cross-platform reporting, systemic remediation
Technologies
AWS, Azure, GCP, Python, Go, Bash, Terraform, CI/CD, LLMs
Responsibilities
Set and scale gold standards for SLOs, on-call readiness, and postmortems; partner with Engineering/Product to align reliability outcomes with roadmaps; design operating cadences for incident reviews and stability planning; automate recurring remediations and build an automation opportunity pipeline; own cross-platform reporting for SLO attainment, MTTR, and cost of failure; apply AI/LLMs to improve triage, incident summarization, and safe automation; lead systemic remediation and major incident response improvements.
Seniority
Senior, hands-on IC with program leadership