CareerPlanSign in

Manager, Site Reliability Engineering

Office - USA - CA - Headquarters💼 Full-time🗓 2026-08-12 → 2026-09-26

Core

Lead and mentor a team of Site Reliability/Production Engineers to ensure the reliability, availability, and operational health of critical Cortex services and infrastructure.

Role type

Manager, Site Reliability Engineering

Builds

Cortex services and infrastructure

Domain

Cybersecurity, Cloud Infrastructure, SaaS

Deliverable

production ML models | infrastructure

Required skills

Site Reliability Engineering, DevOps, Cloud Infrastructure, Kubernetes, containerized environments, large-scale distributed systems, observability, incident management, SLOs, automation, Infrastructure as Code, Python, Terraform, Ansible, GitOps, team leadership

Preferred skills

managing SRE/DevOps teams for large-scale SaaS, multi-region cloud operations, AI-assisted operational capabilities, high-severity incident management

Technologies

Google Cloud Platform (GCP), Prometheus, Thanos, Grafana, OpenTelemetry, Python, Terraform, Ansible, GitOps

Responsibilities

Lead, mentor, and develop a team of Site Reliability/Production Engineers; Own the reliability and operational health of critical Cortex services; Drive improvements in monitoring, alerting, and incident management; Lead major production incidents and root cause analysis; Partner with Engineering teams on service architecture and scalability; Drive automation and self-healing solutions; Establish operational processes and reliability goals; Collaborate on follow-the-sun operations; Evaluate and adopt new technologies for reliability and efficiency.

Seniority

Manager, hands-on leadership

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.