Manager, Site Reliability Engineering
Core
Lead and mentor a team of Site Reliability/Production Engineers to ensure the reliability, availability, and operational health of critical Cortex services and infrastructure.
Role type
Manager, Site Reliability Engineering
Builds
Cortex services and infrastructure
Domain
Cybersecurity, Cloud Infrastructure, SaaS
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering, DevOps, Cloud Infrastructure, Kubernetes, containerized environments, large-scale distributed systems, observability, incident management, SLOs, automation, Infrastructure as Code, Python, Terraform, Ansible, GitOps, team leadership
Preferred skills
managing SRE/DevOps teams for large-scale SaaS, multi-region cloud operations, AI-assisted operational capabilities, high-severity incident management
Technologies
Google Cloud Platform (GCP), Prometheus, Thanos, Grafana, OpenTelemetry, Python, Terraform, Ansible, GitOps
Responsibilities
Lead, mentor, and develop a team of Site Reliability/Production Engineers; Own the reliability and operational health of critical Cortex services; Drive improvements in monitoring, alerting, and incident management; Lead major production incidents and root cause analysis; Partner with Engineering teams on service architecture and scalability; Drive automation and self-healing solutions; Establish operational processes and reliability goals; Collaborate on follow-the-sun operations; Evaluate and adopt new technologies for reliability and efficiency.
Seniority
Manager, hands-on leadership