Lead Software Engineer, Cloud Site Reliability (SRE)
Core
Lead 24x7 cloud reliability operations and incident management for critical Azure-based technology environments.
Role type
Lead Software Engineer, Cloud Site Reliability (SRE)
Builds
Resilient, scalable cloud-native environments on Azure with high availability and SLA adherence
Domain
Cloud Infrastructure / Site Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Azure IaaS, AKS, Kubernetes, Docker, Datadog, Azure Monitor, Terraform, ARM templates, Helm, PowerShell, Python, Bash, ServiceNow, Incident Management, Root Cause Analysis, Operational Reporting
Preferred skills
Multi-cloud (AWS), AIOps, Predictive Monitoring, Anomaly Detection, Self-healing systems, Azure/Datadog/Kubernetes certifications
Technologies
Azure, AKS, Kubernetes, Docker, Datadog, Azure Monitor, Terraform, ARM templates, Helm, Power Automate, ServiceNow, Power BI
Responsibilities
Lead 24x7 NOC operations through rotational shifts; Act as Major Incident Manager for P1/P2 incidents; Manage and troubleshoot Azure infrastructure; Administer AKS, Kubernetes, and Docker environments; Establish observability practices using logs, metrics, and traces; Drive proactive monitoring and AIOps initiatives; Build automation and self-healing workflows; Collaborate on deployment pipelines and cloud-native architecture; Develop operational dashboards and reports; Lead monthly business reviews; Mentor team members and standardize processes
Seniority
Lead, hands-on IC with mentorship responsibilities