Senior Manager, Site Reliability Engineering (Global)
Core
Lead global Site Reliability Engineering (SRE) organization and architect AIOps transformation to enable self-healing, AI-driven observability and incident remediation.
Role type
Senior Manager, Site Reliability Engineering (Global)
Builds
Automated, intelligent platform capabilities for infrastructure products; global SRE practices and 24/7 coverage models.
Domain
Cybersecurity, Cloud Native, AIOps, Machine Learning
Deliverable
production ML models | infrastructure
Required skills
Global team management (15+ engineers), Kubernetes, Cloud Native ecosystems (AWS/GCP/Azure), CI/CD pipelines, ML-driven monitoring (anomaly detection, automated root cause analysis), Financial management of headcount and cloud spend, Policy-as-Code, SLO/Error Budget management, Canary analysis.
Preferred skills
Translating deep tech to business value for C-suite, Influencing R&D leadership on non-functional requirements.
Technologies
Kubernetes, AWS, GCP, Azure, ML/AI frameworks for observability, Policy-as-Code tools.
Responsibilities
Manage and scale multi-geographical SRE team; Standardize global SRE practices and 24/7 coverage; Manage financial aspects of headcount and cloud spend; Drive AIOps roadmap for proactive monitoring; Act as lead consultant for infrastructure product teams; Partner with Platform Engineering to build Golden Paths.
Seniority
Senior Manager, hands-on leadership