About the job
Core
Leading incident response, RCA, and reliability engineering for a company-wide service infrastructure on AWS EKS.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Production infrastructure stability, incident resolution processes, and deployment standards.
Domain
Cloud Infrastructure (AWS), Observability, Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
AWS operations, incident management, RCA, observability platform usage, server/network troubleshooting, system interdependency analysis, chaos engineering planning
Preferred skills
high-traffic service operations, SLI/SLO/error budget management, Java/Spring performance analysis, LLM/agent development and operations, post-mortem culture establishment, deployment automation
Technologies
AWS, EKS, Datadog, Claude Code, Codex
Responsibilities
Lead first-line response and RCA for company-wide service incidents; manage post-mortems and prevent recurrence actions; standardize on-call and incident systems; define SLI/SLOs and quantify reliability; manage production checklists by service level; standardize deployment/change management; plan chaos engineering training
Seniority
Senior, hands-on IC