Senior Manager, SRE
Core
Define and execute SRE strategy to ensure high availability, scalability, and performance of mission-critical systems for Coach and Kate Spade New York.
Role type
Senior Manager, Site Reliability Engineering (SRE)
Builds
Production systems for Coach and Kate Spade New York
Domain
Retail / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
SRE strategy definition, SLO/SLI/Error Budget management, incident management, blameless postmortems, infrastructure automation (IaC), observability stack design, CI/CD pipeline optimization, progressive delivery patterns, team leadership, cross-functional collaboration
Preferred skills
AI/ML and intelligent agents for predictive incident prevention, autonomous remediation, large-scale high-traffic system operations
Technologies
AliCloud, AWS, Kubernetes, Helm, Terraform, Ansible, Pulumi, Prometheus, Grafana, Datadog, Open Telemetry, ELK, Jenkins, GitLab CI, GitHub Actions, ArgoCD, Python, Go, Bash
Responsibilities
Define and execute SRE strategy; Establish and enforce SLOs, SLIs and error budgets; Lead incident management processes; Design and implement automated solutions for provisioning and lifecycle operations; Own and optimize end-to-end CI/CD pipelines; Build, lead, and mentor a growing team of vendor SRE engineers
Seniority
Senior, hands-on IC with leadership responsibilities