Site Reliability and DevOps Engineering Lead
Core
Lead the Platform Reliability & DevOps team to ensure 24x7 high availability, performance, and security for mission-critical clinical decision support systems used by clinicians globally.
Role type
Senior IC Platform Reliability & DevOps Engineering Lead
Builds
Mission-critical clinical platform (drug reference, IV compatibility, pediatric dosing, toxicology databases) serving healthcare organizations
Domain
Healthcare / Clinical Decision Support / Enterprise SaaS
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering, DevOps, CI/CD architecture, distributed system design, database optimization (DB2, Oracle, PostgreSQL), cloud infrastructure, automation scripting (Python, Bash, Java), incident management, capacity planning, vendor management
Preferred skills
Terraform, Docker, Kubernetes, Azure, AI-enabled pipeline optimization
Technologies
DB2, Oracle, Infinispan, OpenLiberty, Azure, Python, Bash, Java, Git, Kubernetes, Docker, Terraform
Responsibilities
Lead and mentor Platform/DevOps engineers; define and enforce platform engineering standards and DevOps practices; own SLIs, SLOs, error budgets, and incident management frameworks; lead Sev1 response and drive systemic fixes; design and scale fault-tolerant distributed systems; standardize and automate end-to-end CI/CD pipelines; lead capacity planning, performance optimization, and cost efficiency; act as technical authority for platform strategy and roadmap.
Seniority
Senior, hands-on IC with strategic leadership