Senior Site Reliability Engineer
Core
Lead reliability strategy and drive AI-powered automation at scale for Zuora's global SaaS platform.
Role type
Senior Site Reliability Engineer (IC with leadership)
Builds
Cloud infrastructure, Kubernetes platforms, and AI-driven automation systems for detection, remediation, and forecasting.
Domain
SaaS / Cloud Infrastructure / AI Operations
Deliverable
production ML models | infrastructure
Required skills
AWS architecture design, Infrastructure-as-Code (Terraform), Python/Shell automation, Linux systems administration, distributed systems operations, Kafka, technical leadership, SLO/SLI definition
Preferred skills
Advanced observability platforms (Prometheus, Grafana, ELK), predictive analytics, anomaly detection, AIOps platforms
Technologies
AWS, Kubernetes, Terraform, Python, Shell, Linux, Kafka, Prometheus, Grafana, ELK
Responsibilities
Define and evolve SLOs, SLIs, and resilience patterns; Build AI-driven automation for detection, remediation, and forecasting; Lead cloud infrastructure and Kubernetes platforms; Drive incident response and operational excellence; Mentor engineers and influence org-wide reliability practices
Seniority
Senior, hands-on IC with leadership