Staff Site Reliability Engineer
Core
Own the reliability, security, and scalability of Lyrebird's production platform serving clinicians during patient consultations.
Role type
Staff Site Reliability Engineer (hands-on technical leadership) (via careerplan.io/jobs/f46807c6-3b33-4bc9-94f2-2db6480fedf9-staff-site-reliability-engineer-at-lyrebird-health)
Builds
AWS infrastructure, deployment pipelines, observability, and incident response systems
Domain
Healthcare technology / Cloud infrastructure
Deliverable
production ML models | infrastructure
Required skills
AWS, Terraform, containers, CI/CD, networking, security, production troubleshooting, observability, SLO definition, incident response, automation, cross-team influence
Preferred skills
ECS/Fargate, PostgreSQL, distributed workloads, TypeScript/Node.js, OpenTelemetry, healthcare technology frameworks (ISO 27001, Cyber Essentials Plus), multi-region operations
Responsibilities
Define SLOs, alerting, and observability strategy; shape incident detection and response; evolve AWS and Terraform platform for resilience and disaster recovery; improve CI/CD pipelines and developer tooling; build security and compliance automation; raise operational maturity across engineering teams; challenge technical decisions regarding reliability and security
Seniority
Staff, hands-on technical leadership