(Seoul) Senior Site Reliability Engineer· Cancer Screening
Core
Ensure stability, availability, and reliability of the Lunit INSIGHT AI product for cancer screening and treatment prediction across diverse global deployment environments.
Role type
Senior Site Reliability Engineer (Cloud Infrastructure)
Builds
Cloud infrastructure, deployment automation, monitoring systems, and incident response processes for the INSIGHT product.
Domain
Healthcare AI / Medical Software / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Azure production environment design, Linux, networking, containers, CI/CD design, incident root-cause analysis, cross-functional collaboration
Preferred skills
SRE practice establishment (on-call, postmortems, SLO/SLA), Infrastructure as Code (Terraform, Bicep), Python/Bash automation, Kubernetes, observability tools (Azure Monitor, Application Insights)
Technologies
Azure, Linux, Docker, Kubernetes, Terraform, Bicep, GitHub Actions, Azure DevOps, Python, Bash, Go, FastAPI, PostgreSQL, Redis, RabbitMQ, Azure Monitor, Application Insights, Log Analytics, PagerDuty, ServiceNow
Responsibilities
Design and enhance cloud infrastructure, deployment, monitoring, and operational automation systems; Diagnose, recover from, and perform root-cause analysis on incidents; Analyze cloud resource usage and drive cost optimization; Collaborate with development teams to build reliability from the design stage; Build and lead a sustainable on-call and incident-response system.
Seniority
Senior, hands-on IC