Vice President, Global Production Operations & Reliability
Core
Lead global production operations, SRE, and DRE teams to ensure the reliability, scalability, and security of a cloud-native SaaS platform for critical event management.
Role type
VP, Global Production Operations & Reliability (Executive)
Builds
Cloud-native SaaS platform for critical event management (Everbridge)
Domain
Cloud Infrastructure / SRE / Platform Engineering
Deliverable
production ML models | product features | infrastructure
Required skills
Global production operations leadership, AWS, Kubernetes, cloud-native architectures, SRE principles, incident management, observability, disaster recovery, change governance, CI/CD, Infrastructure as Code
Preferred skills
Mission-critical industry experience (public safety, healthcare, finance), progressive delivery, service mesh, cloud cost optimization, regulatory frameworks (ISO 27001, SOC 2, NIST, FedRAMP)
Technologies
AWS, Kubernetes
Responsibilities
Define and execute global production operations and reliability strategy; Lead and develop high-performing global SRE and DRE teams; Own operational excellence of AWS and Kubernetes-based cloud platform; Establish best practices for service reliability (SLOs, error budgets, capacity planning); Drive disciplined incident management, change governance, and release management; Lead disaster recovery planning and resilience testing; Partner with Engineering, Product, Security, and Customer Success to embed reliability; Champion automation and continuous improvement; Provide executive leadership and reporting on operational health and strategic initiatives
Seniority
VP, Executive Leadership