Sr. Director, Back-End Engineering
Core
Define and lead company-wide reliability, resilience, scalability, and operational excellence, transforming practices into platform mechanisms and advancing toward an AI-assisted autonomous operating model.
Role type
Sr. Director, Site Reliability Engineering (Head of SRE)
Builds
Platform mechanisms for reliability inheritance, autonomous incident response, disaster recovery, and capacity management for Coupang's global marketplace.
Domain
E-commerce / Distributed Systems / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Organizational leadership, SRE strategy definition, Kubernetes, service mesh, chaos engineering, AI-assisted operations, capacity management, disaster recovery planning, executive stakeholder management, technical architecture decision-making.
Preferred skills
Experience scaling SRE organizations at hyperscalers, building AI-driven incident intelligence, implementing fault injection and game days.
Technologies
Kubernetes, service mesh, cloud-native platforms, AI/ML for operations, observability stacks.
Responsibilities
Set multi-year vision for reliability and autonomous operations; define SRE strategy, standards, and governance; build platform mechanisms for tier-based reliability; lead initiatives to improve availability and resilience; partner with cross-functional leaders to align reliability investments; own executive reliability metrics; build and scale a world-class SRE organization; establish SLOs, SLAs, and error budgets; transform incident management with AI assistance; own disaster recovery and capacity strategy; integrate observability and testing into reliability feedback loops; lead SRE leadership talent and organizational design; guide reliability architecture decisions; deliver measurable improvements in incident metrics and operational efficiency.
Seniority
Senior, hands-on IC with executive leadership scope