Engineering Manager, Site Reliability
Core
Manage a team of senior SRE engineers while personally contributing to building automation, tooling, and operational capabilities for resilient systems.
Role type
Senior Engineering Manager (Site Reliability)
Builds
Automation, tooling, and operational capabilities for production systems
Domain
Financial services / Site Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
People leadership, hands-on software delivery, stakeholder management, coaching senior ICs, roadmap sequencing, risk management, engineering quality standards
Preferred skills
Observability, cloud infrastructure (AWS ECS/Lambda), SRE practices (incident response, chaos engineering), agentic tooling, AI-augmented workflows
Technologies
AWS, ECS, Lambda, infrastructure-as-code
Responsibilities
Lead and coach a team of engineers; personally write code and ship operational tooling; guide solution design and technical decision-making; partner with product and technical leaders to drive adoption; own team outcomes and roadmap sequencing; foster psychological safety and continuous improvement; resolve impediments affecting delivery.
Seniority
Senior, hands-on IC manager