Lead Site Reliability Engineer
Core
Lead investigations into root cause outages, performance, and cost issues for cloud platforms serving Public Safety & Justice markets.
Role type
Senior IC Site Reliability Engineer (Cloud Platform)
Builds
Observable, measurable, reliable, scalable, and maintainable cloud platforms for multi-media evidence management and Emergency Contact Centres.
Domain
Public Safety & Justice / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes (AKS), Azure Cloud, Infrastructure as Code (Bicep/Terraform), Python/PowerShell/C#, Observability (Grafana/Prometheus), Microservices architecture, SLO/SLA management, Database handling (MS-SQL, Elasticsearch)
Preferred skills
Azure DevOps Pipelines, AI tools for automation, ISO 27001/Cyber Essentials/FEDRAMP compliance frameworks
Technologies
Kubernetes, AKS, Azure, Bicep, Terraform, Grafana, Prometheus, Elasticsearch, PowerShell, C#, Python
Responsibilities
Act as gatekeepers of production and manage work backlog; Lead initiatives to automate low-value tasks; Provide technical leadership to Cloud Operations and Support teams; Develop and configure monitoring dashboards and alerts; Optimize system performance, cost, and security through regular reviews and tuning.
Seniority
Senior, hands-on IC