Senior Site Reliability Engineer (Azure)
Core
Lead reliability engineering initiatives, monitoring, and incident response for a large-scale Azure environment to ensure uptime and performance.
Role type
Senior Site Reliability Engineer (Azure)
Builds
Production-ready Azure infrastructure, monitoring dashboards, and automated remediation workflows
Domain
Cloud Infrastructure (Azure) + Observability
Required skills
Azure services (AKS, App Services, Functions, VM Scale Sets, Networking, Security), Terraform/Bicep, CI/CD, PowerShell, Python, ITRS Geneos, SRE practices (SLIs/SLOs, error budgets, chaos testing)
Preferred skills
AZ-400, AZ-305, AZ-104, AZ-500, ITIL v4, SRE Foundation/Practitioner certifications
Technologies
Azure Monitor, Log Analytics, Application Insights, Grafana, Prometheus, OpenTelemetry, ServiceNow, JIRA, GitHub Actions, Azure DevOps
Responsibilities
Lead reliability engineering initiatives and Command Center operations; Design and implement end-to-end monitoring with Azure Monitor, Log Analytics, and ITRS Geneos; Develop alert taxonomies and reduce alert noise; Provide technical support and rapid remediation during P0/P1 incidents; Build automation for alert management and self-healing workflows; Define and uphold SLIs/SLOs and resilience patterns; Enforce desired state using Azure Policy and strengthen security; Partner with engineering and platform teams to improve service reliability.
Seniority
Senior, hands-on IC