SRE / Backend Engineer
Core
Own production reliability, stability, and performance for a clinical AI platform serving 100,000+ physicians across the US and EU, handling billions of tokens and real-time EHR syncs.
Role type
Senior Staff Site Reliability Engineer (SRE) / Backend Engineer
Builds
Scalable, resilient infrastructure and backend services for an ambient clinical documentation AI assistant
Domain
Healthcare technology, Cloud Infrastructure, Machine Learning Operations
Deliverable
production ML models | infrastructure
Required skills
SRE practices, cloud infrastructure (GCP), infrastructure-as-code (Terraform), observability, incident management, API design, system scalability
Preferred skills
Kubernetes, GCP-native tooling (Cloud Run, GKE, Pub/Sub), PostgreSQL at scale, healthcare security/compliance
Technologies
GCP, Terraform, Kubernetes, Cloud Run, GKE, Pub/Sub, PostgreSQL
Responsibilities
Own production reliability and platform stability; scale on-call culture and incident response processes; develop monitoring and observability stacks; drive SRE roadmap including SLOs and chaos engineering; improve developer experience and reduce toil; collaborate with backend, ML, and front-end squads to embed reliability best practices
Seniority
Senior, hands-on IC