Mid SRE – Cloud Product Reliability
Core
Ensure cloud products and services are reliable, observable, scalable, secure, and resilient to protect customers and enable engineering innovation.
Role type
Mid-level Site Reliability Engineer (Cloud Product Reliability)
Builds
Cloud-native applications and services for customers
Domain
Cloud infrastructure and distributed systems
Deliverable
production ML models | product features | infrastructure
Required skills
observability, distributed system architecture, high-availability environment support, cloud governance frameworks, incident management, system performance tuning, reliability tools, Infrastructure as Code (Terraform), monitoring and observability platforms, scripting (PowerShell, Python, Bash), Linux/Windows operations
Preferred skills
designing and implementing reliability engineering practices, troubleshooting, data analysis, technical communication
Technologies
AWS, GCP, Terraform, Linux, Windows, PowerShell, Python, Bash
Responsibilities
Identify reliability risks early and influence architect decisions, establish reliability requirements, validate system behavior under failure conditions, create operational excellence, increase system resilience, drive continuous improvements
Seniority
Mid-level, hands-on IC