Site Reliability Engineer - Mumbai - India
Core
Build and maintain reliability, observability, and operational excellence for .NET and React applications, enabling scalable solutions for clients and staff.
Role type
Site Reliability Engineer (SRE)
Builds
Reliable, scalable, and observable full-stack solutions on Azure
Domain
Professional Services / Audit & Tax / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
.NET (C#), React, Azure platform services, CI/CD pipelines, Infrastructure as Code (Bicep/ARM/Terraform), SRE principles (SLOs/SLIs/error budgets), incident management, root cause analysis, SQL, API testing, code reviews, mentoring
Preferred skills
AIOps platforms, AI-driven observability, AI-assisted tooling (chatbots, runbook copilots), chaos/resilience testing
Technologies
Azure Monitor, Application Insights, Log Analytics, Grafana, GitHub, Azure DevOps, GitHub Actions, Postman, MS SQL
Responsibilities
Define and track SLOs/SLIs/error budgets; design and maintain monitoring/alerting/logging solutions; lead incident response and blameless postmortems; drive adoption of AI-assisted SRE practices; contribute to building internal AI-driven tooling; partner on resilient full-stack solution design; perform system and chaos testing; mentor developers on reliability best practices
Seniority
Mid-Senior, hands-on IC with mentoring responsibilities