Principal Site Reliability Engineer
Core
Design, build, and operate reliable, secure Cloud infrastructure for Palo Alto Networks' large hybrid infrastructure and Wildfire malware protection engine.
Role type
Principal Site Reliability Engineer (Cloud Infrastructure)
Builds
Scalable, high-availability cloud services and automation frameworks for production environments.
Domain
Cybersecurity, Cloud Infrastructure (GCP/AWS), Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Terraform, Ansible, Python, Go, Linux administration, Cloud (GCP/AWS), CI/CD pipelines, Distributed systems troubleshooting, Shell scripting
Preferred skills
GitLab, GitHub, Helm, Bigtable, Memorystore, Bigquery, RabbitMQ, Kafka, MySQL, Spinnaker, Pub/sub
Technologies
Kubernetes, Docker, GCP, AWS, Ansible, Terraform, Vault, Gitlab, Spinnaker, Pub/sub, Bigtable, Memorystore, Bigquery, RabbitMq, Kafka, MySQL, Python, Go
Responsibilities
Design and operate reliable, secure Cloud infrastructure; Develop tools and automation frameworks; Automate robust deployment of services; Orchestrate end-to-end monitoring and alerting; Lead root cause analysis of critical production issues; Mentor and champion SRE culture; Participate in design reviews.
Seniority
Principal, hands-on IC with mentorship