Lead Site Reliability Engineer
Core
Lead a small Site Reliability Engineering (SRE) team responsible for the core production platform, focusing on system reliability, incident response, and operational excellence for a global healthcare AI platform.
Role type
Lead Site Reliability Engineer (IC + Team Lead)
Builds
Production Kubernetes clusters, cloud infrastructure, and core platform services for a global healthcare AI platform serving millions of patients.
Domain
Healthcare technology / Cloud Infrastructure / SRE
Deliverable
production ML models | infrastructure
Required skills
SRE/DevOps leadership, incident response, Kubernetes, cloud infrastructure (AWS), Infrastructure as Code (Terraform), observability (Datadog/Prometheus), SLOs and error budgets, Python/Bash scripting, team hiring and growth.
Preferred skills
Scaling teams through fast headcount growth, regulated/security-sensitive environments, database/queue/cache management.
Responsibilities
Lead end-to-end incident response and on-call rotations; improve operational reliability through automation and process improvements; own and optimize production Kubernetes and cloud infrastructure; build observability solutions (dashboards, alerts, traces); reduce operational toil; manage and grow the SRE team including hiring and career development; define team direction and reliability standards.
Seniority
Senior, hands-on IC with team leadership responsibilities