Principal Site Reliability Engineer
Core
Design, build, and operate reliable, secure Cloud infrastructure for Wildfire, the industry's largest cloud-based malware protection engine.
Role type
Principal Site Reliability Engineer
Builds
Scalable, high-availability cloud infrastructure for cloud-based malware protection services
Domain
Cybersecurity / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Terraform, Ansible, Python, Go, Linux administration, distributed systems troubleshooting, CI/CD pipelines, cloud automation
Preferred skills
GCP expertise, Helm, GitLab, GitHub, RabbitMQ, Kafka, MySQL, BigQuery, Spinnaker, Pub/sub, Bigtable, Memorystore
Technologies
Kubernetes, Docker, GCP, AWS, Ansible, Terraform, Vault, Gitlab, Spinnaker, Pub/sub, Bigtable, Memorystore, Bigquery, RabbitMq, Kafka, MySQL, Python, Go
Responsibilities
Design and operate reliable, secure Cloud infrastructure; Develop tools and automation frameworks; Automate robust deployment of services; Orchestrate end-to-end monitoring and alerting; Lead root cause analysis of critical production issues; Mentor and champion SRE culture; Participate in design reviews
Seniority
Principal, hands-on IC with mentorship