Lead Site Reliability Engineer - Network
Core
Lead SRE responsible for mission-critical network services, reliability culture, and major incident response within a global financial firm.
Role type
Lead Site Reliability Engineer (Network)
Builds
Mission-critical network services, automation pipelines, and reliability workflows for enterprise platforms.
Domain
Financial Services / Network Infrastructure / SRE
Required skills
Site Reliability Engineering (SRE), Network troubleshooting, Python, Java/Spring Boot, .Net, Observability, CI/CD, Container orchestration, Automation (Ansible/Shell), Incident management, Root cause analysis, Service Level Objectives (SLOs), Risk analysis (FMEA), AI-assisted workflow validation
Preferred skills
Large-scale network operations, Leadership in major incidents, Advanced automation track record, SRE concepts mastery, Stakeholder management, Self-education in new technologies
Technologies
Python, Java, Spring Boot, ASP.NET, Ansible, Shell, AI tools, SD-WAN, SDA, SDN, Firewalls, Load balancers, Proxies
Responsibilities
Model and promote SRE culture and practices; Lead initiatives to improve reliability and stability of applications and platforms; Serve as primary contact during major incidents; Own reliability for mission-critical network services; Design and deliver automation to reduce toil; Provide technical leadership across network domains (SD-WAN, SDA, SDN, routing, switching, security); Lead reuse-first adoption of AI-assisted reliability workflows.
Seniority
Lead, hands-on IC with mentorship responsibilities