CareerPlanGet AI match score →

Senior Site Reliability Engineer

San Francisco or Palo Alto, CA💼 Full-time🗓 2026-05-28 → 2026-07-31

Core

Architecting and operating global production systems for a distributed computing platform serving software developers and data scientists.

Role type

Senior Site Reliability Engineer (Infrastructure Strategy)

Builds

Autonomous, self-healing infrastructure for Ray (distributed computing platform)

Domain

Cloud infrastructure, distributed systems, machine learning platforms

Deliverable

infrastructure

Required skills

distributed systems architecture, multi-cloud management, Kubernetes, infrastructure as code, observability, incident management, SLO definition, technical mentorship

Preferred skills

Python or Go programming, large-scale microservices experience

Technologies

AWS, GCP, Azure, Terraform, Kubernetes

Responsibilities

Architect unified cloud component utilization strategies, design autonomous self-healing infrastructure, build robust observability systems (metrics, logging, tracing), establish testing infrastructure, define organization-wide SLOs and error budgets, implement on-call and incident management systems, coordinate cloud service deployments.

Seniority

Senior, hands-on IC with mentorship responsibilities

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗