SRE, Azure Cloud
Core
Senior SRE responsible for managing Stage and Production environments, ensuring reliability and availability of customer-facing Azure-based microservices.
Role type
Senior hands-on IC Site Reliability Engineer
Builds
Production-ready infrastructure, monitoring solutions, and operational runbooks for enterprise SaaS applications
Domain
Cloud operations and SRE for Azure-based microservices
Required skills
Azure services (AKS, VMs, App Services, Storage, Networking, Key Vault), Kubernetes, Docker, Helm, CI/CD (Azure DevOps, Bitbucket, Git), observability (Grafana, Prometheus, Graylog, Azure Monitor, Application Insights), database operations (SQL Server, PostgreSQL, MySQL, Cosmos DB, Redis), scripting (PowerShell, Python, Bash), incident management, root cause analysis
Preferred skills
Enterprise SaaS support, 24x7 on-call operations, SRE practices (SLIs, SLOs, MTTR, service availability management)
Technologies
Azure, AKS, Kubernetes, Docker, Helm, Azure DevOps, Bitbucket, Git, Grafana, Prometheus, Graylog, Azure Monitor, Application Insights, Log Analytics, Cosmos DB, SQL Server, PostgreSQL, MySQL, Redis, PowerShell, Python, Bash
Responsibilities
Manage and support Stage and Production environments; deploy applications, infrastructure, and databases; build and maintain monitoring solutions and alerts; diagnose production incidents and performance issues; lead incident response and conduct root cause analyses; create operational runbooks; partner with Engineering, Product, and Customer Support teams during major incidents (via careerplan.io/jobs/Sierra-Business-Solution-SRE-Azure-Cloud-sre-azure-cloud-at-sierra-business-solution)
Seniority
Senior, hands-on IC