Senior Site Reliability Engineer
Core
Ensuring the reliability, scalability, performance, and operational excellence of Merlin.net, a remote monitoring platform for patients with implanted cardiac devices.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Resilient, fault-tolerant systems and infrastructure for a remote monitoring platform serving cardiologists and care teams.
Domain
Healthcare / Medical Devices / Cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Systems/automation programming (Python, Go, Bash, PowerShell), Microsoft Azure (AKS, Monitor, DevOps), Kubernetes/Docker orchestration, Observability (Prometheus, Grafana, ELK, Datadog), CI/CD pipeline design, Distributed systems architecture, Linux & networking fundamentals, Incident management & RCA
Preferred skills
Regulated healthcare/medical device environment experience, HIPAA compliance familiarity, Cloud cost optimization (FinOps), Disaster recovery design
Technologies
Azure, Kubernetes, Docker, Prometheus, Grafana, ELK, Datadog, Azure Monitor, Azure DevOps
Responsibilities
Design and maintain highly available, fault-tolerant systems; Identify and eliminate performance bottlenecks; Define and monitor SLIs/SLOs; Develop monitoring, logging, and alerting solutions; Automate operational tasks; Scale services and infrastructure; Collaborate with engineering and security teams; Create documentation and runbooks; Lead blameless postmortems.
Seniority
Senior, hands-on IC