Senior Site Reliability Engineer
Core
Operate and maintain a large-scale GCP environment, design and enhance observability systems, and ensure the reliability of the Cortex SecOps platform (XDR, XSIAM, XSOAR, XPANSE).
Role type
Senior Staff Site Reliability Engineer (Observability)
Builds
Production observability infrastructure and monitoring solutions for the Cortex SecOps platform
Domain
Cybersecurity / Cloud Infrastructure (GCP)
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
GCP, Prometheus, Grafana, Open Telemetry, Pagerduty, Kubernetes, Docker, Python, Linux Shell, Ansible, Terraform
Preferred skills
None stated
Technologies
GCP, Prometheus, Grafana, Open Telemetry, Pagerduty, Kubernetes, Docker, Ansible, Terraform
Responsibilities
Optimize cloud infrastructure using cloud-native technologies; improve monitoring processes, alerts, and metrics; manage incidents to ensure efficient resolution; automate complex monitoring and alerting tasks; provide follow-the-sun on-call coverage; collaborate with engineering teams to influence product operability.
Seniority
Senior Staff, hands-on IC