ENGINEERING MANAGER - SRE
Core
Lead the Site Reliability Engineering (SRE) team to own and drive the reliability programme, focusing on high availability and disaster recovery (HA/DR) testing and incident management across Asia-Pacific and EMEA regions.
Role type
Engineering Manager for Site Reliability Engineering (SRE)
Builds
Reliability programme, HA/DR testing cadence, incident command protocols, and a high-performing SRE team
Domain
Cloud Infrastructure & Site Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
SRE or infrastructure leadership, incident command at scale, HA/DR programme building, software engineering, AI-assisted tooling fluency, stakeholder management, distributed team management, mentorship
Preferred skills
None explicitly stated
Technologies
None explicitly stated
Responsibilities
Build and mature continuous HA/DR testing cadence, establish and optimize incident command protocols, close cross-team reliability gaps, grow the SRE team from two to four engineers
Seniority
Manager, hands-on leadership