Lead Site Reliability Engineer (SRE)
Core
Lead Site Reliability Engineer ensuring platform stability, operational resilience, and continuous improvement for global business operations.
Role type
Lead Site Reliability Engineer (SRE)
Builds
Resilient products and scalable distributed systems for global business operations
Domain
Financial services / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, containerization (Docker, ACR), public cloud platforms (AWS, Azure, GCP), DevOps practices, configuration management, automation, observability, security implementations, certificate management, encryption methodologies, distributed system design, troubleshooting, scripting, pipeline management, software design
Preferred skills
Azure DevOps (AZ-400), Azure Cloud Developer (AZ-203)
Technologies
AWS, Azure, GCP, Docker, Kubernetes, Azure DevOps, GitHub, Matrix
Responsibilities
Ensure platform stability and health, support developers in producing resilient products, assist in application build phases with operational design and monitoring, foster agile culture and enforce operational standards, engage in development lifecycle to optimize customer experience, monitor system behavior and detect anomalies, conduct blameless post-mortems
Seniority
Lead, hands-on IC with mentoring capabilities