Lead Site Reliability Engineer - SRE/DevOps
Core
Lead a team of Site Reliability Engineers to establish and execute a comprehensive reliability strategy for the trading portfolio, focusing on infrastructure, network connectivity, application performance, and throughput.
Role type
Senior IC/Lead Site Reliability Engineer (SRE)
Builds
High-reliability trading infrastructure and services on Microsoft Azure
Domain
Financial services / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering, DevOps, SLA/SLO/SLI management, Microsoft Azure (IaaS/PaaS), Dynatrace, observability, incident management, team leadership, automation
Preferred skills
APM technologies, high-stakes operational environments
Technologies
Azure DevOps, Dynatrace, IaaS, PaaS, Sitecore
Responsibilities
Direct and manage a team of Site Reliability Engineers, offering technical guidance and mentorship; Own and refine the SLA/SLO/SLI framework including error budgets; Set up and enhance monitoring and alerting systems across infrastructure and applications; Lead major incident management efforts and conduct blameless postmortems; Assess application and infrastructure performance to identify root causes; Participate in 24/7 support rotation; Recognize automation opportunities to bolster reliability and efficiency.