Sr. Manager, Incident Management and Site Reliability Engineering
Core
Lead SRE team to ensure resilience and observability of internal business systems (Finance, HR, Supply Chain) and global SaaS ecosystem.
Role type
Senior Manager, Site Reliability Engineering
Builds
Resilient internal SaaS ecosystem (NetSuite, Coupa, Workday) and underlying network infrastructure
Domain
Enterprise SaaS, Business Continuity, Internal Systems
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE/DevOps leadership, Order-to-Cash/Procure-to-Pay lifecycle expertise, Enterprise ecosystem management, Networking (SD-WAN/VPNs), Identity management, Observability tools, Infrastructure as Code
Preferred skills
Python, Go, Terraform, Stakeholder communication
Technologies
NetSuite, Coupa, Workday, Datadog, Splunk, New Relic, Prometheus, Okta, Azure AD, Terraform, Python, Go
Responsibilities
Lead and mentor SRE team, Architect observability across business paths, Define and track SLOs/Error Budgets, Own Major Incident Response process, Lead Root Cause Analysis, Oversee API-driven connections and identity management
Seniority
Senior, hands-on IC with people management