Manager, Site Reliability Engineering
Core
Lead the reliability, stability, and operational excellence of enterprise platforms, owning 24x7 incident management and SRE engineering efforts for high system availability.
Role type
Senior IC SRE Manager
Builds
Production systems, AI/ML workloads, and automated operational solutions for global brands
Domain
Cloud Infrastructure & Site Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
SRE leadership, incident management, root cause analysis, observability, automation, AWS, Kubernetes, microservices, release management, SLO/SLI design, cloud cost optimization
Preferred skills
AI/ML workload support, security best practices, team mentoring, strategic alignment
Technologies
AWS, Kubernetes (EKS), Datadog, CloudWatch, ELK, Prometheus, Grafana
Responsibilities
Own end-to-end reliability of production systems within SLAs; Lead and govern a 24x7x365 incident management team; Drive blameless RCA culture and action item closure; Improve observability and reduce alert noise; Drive automation to reduce operational toil; Mentor a team of ~14 engineers; Optimize cloud usage and reduce spend.
Seniority
Senior, hands-on IC with people leadership