Senior Site Reliability Engineer
Core
Own the reliability, performance, and operational excellence of cloud-native software platforms by bridging development and operations.
Role type
Senior Site Reliability Engineer
Builds
Cloud-native software platforms, internal developer platforms, and automation tooling.
Domain
Cloud Infrastructure (AWS, GCP, Azure)
Deliverable
production ML models | infrastructure
Required skills
Incident response and on-call management, Infrastructure as Code (Terraform, Ansible), Kubernetes cluster management, CI/CD pipeline management, Capacity planning and load testing, SLO/SLI definition and monitoring, Root cause analysis (RCA), Grafana dashboarding
Preferred skills
Multi-cloud architecture expertise, Cost optimization strategies, Platform engineering practices
Technologies
Terraform, Ansible, AWS, GCP, Azure, Kubernetes, GitHub Actions, Grafana
Responsibilities
Lead incident response and on-call management for production outages; Design and maintain infrastructure-as-code (IaC) for multi-cloud environments; Manage Kubernetes cluster lifecycle and security; Build and optimize CI/CD pipelines; Lead capacity planning and performance benchmarking; Develop internal developer platforms to improve engineering velocity.
Seniority
Senior, hands-on IC