Manager, Site Reliability Engineering (Auth0)
Core
Lead the Site Reliability Engineering team to ensure the reliability, scalability, and resilience of Auth0, a trusted authentication platform for millions of users worldwide.
Role type
Manager, Site Reliability Engineering (IC + Leadership)
Builds
Scalable, resilient cloud-native infrastructure for Auth0
Domain
Identity & Access Management (IAM), Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Cloud-native architecture, Infrastructure as Code (Terraform), Container orchestration (Kubernetes), Microservices, Go or Python programming, SRE principles, Team leadership, Incident management, Observability design
Preferred skills
Open-source contributions, Runbook automation, Strategic technical planning
Technologies
AWS, Azure, Terraform, Kubernetes, Go, Python
Responsibilities
Define technical direction and roadmaps for the SRE team, participate in 24/7 on-call rotations to troubleshoot critical incidents, design and implement monitoring and automation to reduce toil, establish reliability policies and cultural standards, mentor and develop SRE talent, represent reliability in architectural reviews and strategic planning
Seniority
Manager, hands-on IC with team leadership