Senior Reliability Engineer
Rewrite
## Senior Reliability Engineer
We are looking for a Senior Site Reliability Engineer to join our team delivering and supporting critical applications running on Azure. The ideal candidate will be an expert in Azure services, have a combination of SRE and DevOps skills including automation, monitoring, observability, CI/CD, incident management, and have a deep understanding of end to end application workflow. As a Senior Site Reliability Engineer, you will play a pivotal role in ensuring the reliability and performance of our applications throughout the lifecycle.
## Responsibilities
- Investigates and resolves complex incidents escalated to the team. Runs post incident review sessions and implements fixes and improvements.
- Conduct service transition activities including establishing metrics to track performance, setting up monitoring, Runbook updates, executing Game Day/OAT, and support team training.
- Maintains services once they are live by measuring and monitoring availability, latency, and overall system health.
- Scales systems sustainably through mechanisms like automation and observability, evolving systems by advocating for changes that improve reliability and velocity.
- Maintains scalable and efficient CI/CD pipelines for application enhancement and fixes.
- Conducts regular capacity and finops review based on usage trends and growth projections.
- Develops disaster recovery (DR) plans and conducts regular DR testing to validate recovery procedures and identify areas for improvement.
- Ensures application compliance with regulatory and security requirements.
- Proactively continues to build and apply relevant domain knowledge that may relate to workflows, data pipelines, business policies, configurations, and constraints.
- Coordinates on security principal access management and triages security issues.
## Requirements
- Degree in Computer Science, Software Engineering, Electronics/Electrical Engineering, or equivalent.
- 5+ years of experience working as a site reliability engineer or DevOps engineer responsible for application availability and reliability, implementing automation, and optimizing system performance.
- Extensive hands-on experience with Azure services preferably Microsoft Fabric and Purview.
- Familiarity with infrastructure-as-a-code tools such as Terraform and Azure Resource Manager.
- Scripting and automation skills using Python, PowerShell, or other languages.
- Strong knowledge of ITIL framework and best practices for incident, change, configuration, and problem management.
- Have a good understanding of REST API.
- Excellent English communication skill. Must be able to work with stakeholders located globally.
- Excellent troubleshooting skills and ability to analyze complex issues.
## Nice to Have
- Intellectually curious people, passionate about the bigger picture of how technology industry is evolving, ready to ask difficult questions and deal with complicated scenarios.
- Creative and a problem solver.
## Benefits
- Personal development through a wide range of learning tools both formal and informal.
- Healthcare, retirement planning, paid volunteering days and wellbeing initiatives.
- Equal opportunities employer.
- Reasonable accommodation for religious practices, mental health or physical disability needs.
- Great Place to Work certified in India.
- Join us and be part of a team that values innovation, quality, and continuous improvement.
Sourced via efinancialcareers · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.