Site Reliability Engineer - SaaSOps
Core
Define and embed SRE best practices, manage incident response, and ensure high availability for a SaaS platform serving life sciences companies.
Role type
Senior Site Reliability Engineer (SaaS)
Builds
High-availability SaaS platform with hybrid cloud/on-prem deployments
Domain
Life Sciences / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE best practices, SLA/SLI/SLO definition, incident management, root cause analysis, automation scripting, hybrid cloud architecture, tenant isolation, observability framework design, CI/CD pipeline optimization, GitOps
Preferred skills
PostgreSQL performance tuning, Kubernetes internals, Terraform, Ansible, Azure Monitor, Application Insights, Log Analytics, Prometheus, Grafana
Technologies
Python, PowerShell, Microsoft Azure, Terraform, Ansible, Kubernetes, PostgreSQL, Azure Monitor, Application Insights, Log Analytics, Prometheus, Grafana
Responsibilities
Define and embed SRE best practices, establish SLA/SLI/SLOs, design disaster recovery strategies, automate manual processes, manage incident response, ensure tenant isolation, strengthen system resiliency, lead incident response efforts, drive root cause analysis, develop operational runbooks, design observability framework, troubleshoot CI/CD pipelines
Seniority
Senior, hands-on IC