IN_Senior Associate_Azure SRE Devops Kubenetes_GCC_Advisory_Bangalore
Core
Site Reliability Engineer ensuring high availability, reliability, and operational excellence for large-scale, business-critical systems in a Global Capability Center.
Role type
Senior IC Site Reliability Engineer (SRE)
Builds
Highly scalable, global platforms and cloud environments
Domain
Cloud Infrastructure & Site Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work -> infrastructure
Required skills
Linux/Unix systems, Python/Go/Java/Bash, Cloud platforms (Azure/AWS/GCP), Kubernetes & Docker, CI/CD pipelines, Monitoring & observability tools, Infrastructure as Code
Preferred skills
24x7 production support experience, SRE best practices (SLIs/SLOs/error budgets), FinOps/cost optimization
Technologies
Azure, AWS, GCP, Kubernetes, Docker, Terraform, ARM, CloudFormation, Pulumi, Prometheus, Grafana, Datadog, ELK, Azure Monitor
Responsibilities
Ensure high availability and reliability of large-scale systems, Monitor, troubleshoot, and resolve production incidents, Define and track SLIs/SLOs and error budgets, Perform root cause analysis and drive preventive actions, Build and maintain automation scripts and runbooks, Support and optimize cloud environments, Improve system resilience through capacity planning and failover testing
Seniority
Senior, hands-on IC