Site Reliability Engineer Leader
Core
Lead Site Reliability Engineering efforts to ensure reliability, resiliency, and innovation in mission-critical information systems and ecosystems for enterprise customers.
Role type
Senior IC Site Reliability Engineer (SRE) with leadership responsibilities
Builds
End-to-end services spanning customer sites, platforms, and hyperscaler clouds
Domain
Enterprise IT operations, Cloud Infrastructure, and System Reliability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Operational management, Incident management, Application monitoring, Service level objectives (SLOs), Enterprise server administration (Windows/Linux/UNIX), Hyperscaler Cloud platforms, Scripting (JSON/YAML/Bash/PowerShell)
Preferred skills
Infrastructure as Code (Ansible/Terraform), Python, Kubernetes, Open-source observability tools (Prometheus/Grafana/Loki)
Technologies
AWS, Azure, Google Cloud Platform, OpenShift, RHEL, AIX, Solaris, Kubernetes, Prometheus, Grafana, Loki, Ansible, Terraform, Python, Bash, PowerShell
Responsibilities
Analyze business needs and provide strategic advice and designs for system reliability, Implement strategies to cap operations load and handle overflow using appropriate tooling and metrics, Collaborate with stakeholders, business, development, and DevSecOps teams to define service level indicators and objectives, Work on end-to-end services spanning customer sites and platforms, Identify and mitigate common operational issues to ensure robustness and security, Partner with customers to build trusted relationships and drive growth
Seniority
Senior, hands-on IC with strategic leadership