Service Excellence Engineer
Core
Lead technical response during critical service incidents, ensuring swift recovery and minimal business disruption across warehouses, offices, and GSC environments.
Role type
Senior IC Service Excellence Engineer (Incident Response & Operational Resilience)
Builds
Early-warning systems, real-time visibility dashboards, and automation-driven remediation workflows for critical services.
Domain
Logistics/Supply Chain (SbM-supported services) + Cloud Infrastructure & Observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Incident response leadership, root-cause analysis, observability platform usage, dashboard development, automation scripting, cross-functional collaboration, continuity backlog management, runbook creation, metrics improvement.
Preferred skills
Multi-region operational environment experience, cloud platform expertise, network infrastructure knowledge.
Technologies
Observability platforms, monitoring systems, event analysis tools, cloud platforms, network infrastructure, automation scripts.
Responsibilities
Lead technical response during critical service incidents, build early-warning and real-time visibility using observability platforms, conduct structured root-cause analysis and drive permanent corrective actions, implement automation-driven remediation steps, maintain and prioritise a continuity backlog, create runbooks and playbooks for consistent execution, provide continuity insights and incident learnings to leadership.
Seniority
Senior, hands-on IC