Sr. Database Site Reliability Engineer (DB SRE)
Core
Own the reliability, availability, and operational maturity of business-critical Azure PostgreSQL platforms supporting the CoverMyMeds platform.
Role type
Senior individual contributor Database Site Reliability Engineer
Builds
Cloud database infrastructure and high-availability, disaster recovery strategies for stateful systems
Domain
Healthcare technology, Cloud Infrastructure (Azure)
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
PostgreSQL operations, Infrastructure as Code (Terraform), cloud security principles, database observability, incident response, performance tuning, replication, backup/restore, disaster recovery
Preferred skills
Kubernetes (AKS), Helm, ArgoCD, CI/CD pipelines, Git/GitOps workflows, regulated environment experience
Technologies
Azure, PostgreSQL, Terraform, Datadog, Kubernetes, AKS, Helm, ArgoCD
Responsibilities
Design and operate cloud database infrastructure using Infrastructure as Code, lead incident response for database-related production issues, define and validate high availability and disaster recovery strategies, troubleshoot complex issues across performance and replication, enforce least-privilege access and security compliance, provide senior technical leadership and mentor less-experienced engineers
Seniority
Senior, hands-on IC with technical leadership