Director, Site Reliability Engineering
Core
Lead a global Site Reliability Engineering team to ensure the reliability, scalability, and cost-efficiency of Counterpart's multi-tenant healthcare infrastructure.
Role type
Senior Manager, Site Reliability Engineering
Builds
Multi-tenant infrastructure for primary care physicians and patient management tools
Domain
Healthcare technology / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Team leadership and hiring, Kubernetes, GCP (GKE, Cloud SQL, Pub/Sub, GCS), Terraform, Helm, ArgoCD, PostgreSQL, Prometheus/Grafana, Python, Go, CI/CD pipelines, developer tooling automation, FinOps, SLO/SLI definition, AI tooling integration
Preferred skills
Experience leading teams across multiple time zones, building internal tooling, sound build vs. buy judgment
Technologies
Kubernetes, GCP, Terraform, Helm, ArgoCD, PostgreSQL, Prometheus, Grafana, GitHub Actions, Claude Code
Responsibilities
Lead and grow the SRE team across multiple time zones, build strategic partnerships with product engineering to shift from reactive to proactive reliability, scale multi-tenant infrastructure for new customer onboarding, own cloud cost management and FinOps practices, champion developer self-service and platform engineering, ensure the SRE team leverages AI tooling in workflows
Seniority
Senior Manager, people leadership & technical strategy