Principal Site Reliability Engineer
Core
Champion end-to-end resiliency, incident management, and operational excellence for mission-critical commerce infrastructure at scale.
Role type
Principal Site Reliability Engineer (Resiliency & Operational Excellence)
Builds
Incident management processes, real-time system visibility dashboards, and organization-wide peak-event readiness programs.
Domain
E-commerce / Cloud Infrastructure / Site Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Incident management, operational excellence, system scaling for peak loads, data analysis for process improvement, cross-functional leadership, project management, technical communication
Preferred skills
Experience running technical training, mentoring others, learning new technologies
Technologies
(Not explicitly listed)
Responsibilities
Build and champion end-to-end incident detection, response, and postmortem processes; Develop real-time metrics and dashboards to track system health; Use operational data to identify gaps and partner with engineering teams on fixes; Lead organization-wide readiness programs for high-traffic events like Black Friday; Identify resilience gaps and collaborate on solutions across infrastructure and product teams; Drive organizational communication, documentation, and training on resiliency.
Seniority
Principal, strategy & mentorship
