Engineering Manager - Site Reliability & Observability (x/f/m)
Core
Lead an SRE team to ensure platform reliability, scalability, and resilience across infrastructure and observability for a large-scale healthcare platform.
Role type
Engineering Manager, Site Reliability Engineering
Builds
Reliable, scalable, and resilient platform infrastructure supporting 170+ applications, 520,000 health professionals, and 90 million patients.
Domain
Healthcare technology, Cloud-native infrastructure, Observability
Deliverable
production ML models | infrastructure
Required skills
People leadership, Cloud-native environments (AWS, GCP, Kubernetes), Observability tooling and architecture, Infrastructure as code (Terraform), Secrets management, Incident response, Postmortem analysis, Cross-functional collaboration
Preferred skills
Scaling SRE/platform teams in high-traffic environments, Backend programming (Go, Python, Ruby), Driving cultural/technical transformations, Experience in regulated environments (healthcare, fintech)
Technologies
Kubernetes, Terraform, AWS, GCP, Prometheus, OpenTelemetry, Datadog, Vault, Elasticsearch
Responsibilities
Lead, coach, and grow a team of Site Reliability Engineers; Define and evolve reliability and observability strategy; Own the team's on-call experience and incident response processes; Collaborate with product and engineering teams to align reliability capabilities with platform needs.
Seniority
Manager, people leadership with technical depth