Director, Site Reliability Engineering (Production Engineering)
Core
Lead SRE teams to ensure 99.99%+ availability for a cloud platform processing hundreds of billions of daily transactions, owning system architecture and automated incident mitigation.
Role type
Director, Site Reliability Engineering (Production Engineering)
Builds
Cloud-native platforms running on Kubernetes and multi-cloud infrastructure
Domain
Cloud Security / SASE / Zero Trust
Deliverable
production ML models (via careerplan.io/jobs/5237915007-director-site-reliability-engineering-production-engineering-at-zscaler)
Required skills
Site reliability engineering principles, chaos testing, capacity planning, automated incident triage, distributed tracing, high-cardinality metrics engines, distributed log indexing, streaming telemetry pipelines, Kubernetes, multi-cloud architecture, managing distributed engineering teams
Preferred skills
AI/ML technologies and workflows, high-volume data ingestion, backpressure management, real-time aggregation, data lifecycle management, executive-level communication
Technologies
Kafka, OpenTelemetry, Prometheus, ClickHouse, Elasticsearch
Responsibilities
Take end-to-end operational accountability for Tier-0 and Tier-1 platforms, Enforce multi-nines SLAs and SLOs, Oversee capacity planning and resource forecasting, Implement high-signal alerting frameworks, Lead Production Engineering across India
Seniority
Director, hands-on IC with management