Principal Site Reliability Engineer (ADEM)
Core
Principal SRE architecting self-healing infrastructure, observability, and automation for Palo Alto Networks' ADEM platform to ensure high-scale synthetic and real-user monitoring reliability.
Role type
Principal Site Reliability Engineer (IC)
Builds
Cloud-native infrastructure, CI/CD pipelines, and automated observability tools for end-to-end digital experience management.
Domain
Cybersecurity / Cloud Infrastructure / Observability
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Infrastructure as Code (Terraform), Cloud platforms (GCP/AWS), Python/Go, GitOps, Root Cause Analysis, AI productivity tools
Preferred skills
Kafka, Policy-as-code, Data streaming frameworks
Technologies
Terraform, Kubernetes, GitLab CI, ArgoCD, Prometheus, Grafana, Loki, Docker, GCP, AWS, Vault, Kafka, MySQL, Python, Bash, Go
Responsibilities
Architect "Golden Paths" for service delivery with integrated SLOs and automated canary analysis; Design and operate reliable, secure Cloud infrastructure for high-scale monitoring; Lead root cause analysis of critical production issues; Develop automation frameworks for IaC and Monitoring as Code; Drive CI/CD and AIOps initiatives.
Seniority
Principal, hands-on IC