Manager, Site Reliability Engineering
Core
Lead a Site Reliability Engineering team to ensure the reliability, scalability, and security of Litera's cloud and on-premises platforms for top law firms.
Role type
Manager, Site Reliability Engineering
Builds
Cloud and on-premises SaaS platforms for legal professionals
Domain
Legal Technology / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Distributed SaaS operations, SRE team leadership, major incident management, observability and monitoring, infrastructure automation, CI/CD practices
Preferred skills
Azure or AWS cloud platforms, ITIL/Problem Management, regulated environment experience
Technologies
Terraform, Ansible, Azure, AWS
Responsibilities
Lead and develop a high-performing SRE team, improve platform reliability and performance, drive incident response and root cause analysis, establish reliability metrics and dashboards, reduce operational toil through automation, partner with Engineering to improve application reliability, support platform growth and cost optimization
Seniority
Manager, hands-on leadership