Staff Site Reliability Engineer
Core
Ensure availability, performance, and operational maturity of cloud-native platforms and critical data/AI/ML services.
Role type
Staff Site Reliability Engineer (Senior IC)
Builds
Cloud-native platform, data services, AI/ML services
Domain
Life sciences, diagnostics, biotechnology
Deliverable
production ML models | infrastructure
Required skills
SRE principles (SLOs, error budgets, toil reduction), observability platform design, cloud platform expertise, container orchestration, Infrastructure as Code, automation scripting, distributed systems troubleshooting, cross-team architecture leadership, technical mentorship
Preferred skills
Life sciences domain experience, chaos engineering, resilience testing, capacity planning, FinOps
Technologies
Kubernetes, Docker, Terraform, OpenTofu, Pulumi, Prometheus, Grafana, Datadog, ELK stack, OpenTelemetry, Splunk, AWS, Azure, GCP, Python, Go
Responsibilities
Champion SRE practice at scale by establishing and monitoring Service Level Objectives and error budgets; Own observability end-to-end by designing and maintaining monitoring, logging, and distributed tracing; Eliminate toil by automating repetitive operational work; Lead incident response and blameless postmortems; Shape systems before they are built by partnering with development teams; Set technical direction and grow the team through standards, documentation, and mentorship
Seniority
Staff, hands-on IC with strategic influence