CareerPlanSign in

About the job

Korea💼 Full-time🗓 2026-09-22 → 2026-09-25

Core

Leading incident response, RCA, and reliability engineering for a company-wide service infrastructure on AWS EKS.

Role type

Senior Site Reliability Engineer (SRE)

Builds

Production infrastructure stability, incident resolution processes, and deployment standards.

Domain

Cloud Infrastructure (AWS), Observability, Reliability Engineering

Deliverable

production ML models | infrastructure

Required skills

AWS operations, incident management, RCA, observability platform usage, server/network troubleshooting, system interdependency analysis, chaos engineering planning

Preferred skills

high-traffic service operations, SLI/SLO/error budget management, Java/Spring performance analysis, LLM/agent development and operations, post-mortem culture establishment, deployment automation

Technologies

AWS, EKS, Datadog, Claude Code, Codex

Responsibilities

Lead first-line response and RCA for company-wide service incidents; manage post-mortems and prevent recurrence actions; standardize on-call and incident systems; define SLI/SLOs and quantify reliability; manage production checklists by service level; standardize deployment/change management; plan chaos engineering training

Seniority

Senior, hands-on IC

Sourced via skcareers · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.