Senior Platform SRE
Core
Ensure reliability, operability, and continuous improvement of enterprise platforms across hybrid cloud and on-prem environments through engineering-driven operations.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Operational automation, IaC configurations, observability standards, and AIOps enablement for enterprise platforms.
Domain
Enterprise IT Operations, Hybrid Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
L3 incident leadership, Infrastructure-as-Code (Terraform/Ansible), Python scripting, hybrid cloud fundamentals, ITSM processes, observability design, RCA and problem management
Preferred skills
AIOps/anomaly detection, containerization, CI/CD pipelines, virtualization and backup/DR, ML/DL for operational analytics
Technologies
Terraform, Ansible, Python, PowerShell, Bash, Azure, ITSM tools
Responsibilities
Define SLOs/KPIs and lead operability gates; design/build operational automation and IaC; lead diagnosis and recovery for major incidents; define actionable signals and alert quality; advance predictive operations and AIOps; equip outsourced L1/L2 providers with runbooks and governance; partner with Platform Engineering to ensure operable-by-design capabilities
Seniority
Senior, hands-on IC with mentorship