Senior Manager - Site Reliability Engineering (SRE)
Core
Lead a team of Site Reliability Engineers to ensure the reliability, performance, and availability of Linux-based systems and infrastructure powering digital commerce platforms.
Role type
Senior Manager, Site Reliability Engineering (Leadership + Hands-on IC)
Builds
Digital commerce platform stability, automation, and observability
Domain
Enterprise IT Infrastructure / Digital Commerce
Deliverable
production ML models | product features | infrastructure
Required skills
Linux system administration, Kubernetes, Docker, CI/CD pipeline design, Infrastructure as Code (Terraform), Configuration Management (Puppet), Python scripting, Load balancing, Virtualization, Observability, Incident management, SRE principles (SLIs/SLOs)
Preferred skills
Mentoring senior engineers, Organizational change management, Cross-functional collaboration
Technologies
Kubernetes, Docker, GitHub Actions, Terraform, Puppet, Python, F5, VMware, Nginx, Apache Tomcat, Datadog, JFrog Artifactory
Responsibilities
Lead and develop a high-performing SRE team; Set technical and strategic vision for the SRE domain; Establish and improve CI/CD standards; Define configuration management and automation standards; Guide strategy for load balancing and virtualized infrastructure; Champion observability practices; Own incident management and root cause analysis; Drive adoption of SRE/DevOps principles; Recruit, mentor, and retain top SRE talent; Partner with Product, Architecture, and Security teams.
Seniority
Senior, hands-on IC with management responsibilities