Senior Site Reliability Engineer
Core
Architecting and operating global production systems for a distributed computing platform serving software developers and data scientists.
Role type
Senior Site Reliability Engineer (Infrastructure Strategy)
Builds
Autonomous, self-healing infrastructure for Ray (distributed computing platform)
Domain
Cloud infrastructure, distributed systems, machine learning platforms
Deliverable
infrastructure
Required skills
distributed systems architecture, multi-cloud management, Kubernetes, infrastructure as code, observability, incident management, SLO definition, technical mentorship
Preferred skills
Python or Go programming, large-scale microservices experience
Technologies
AWS, GCP, Azure, Terraform, Kubernetes
Responsibilities
Architect unified cloud component utilization strategies, design autonomous self-healing infrastructure, build robust observability systems (metrics, logging, tracing), establish testing infrastructure, define organization-wide SLOs and error budgets, implement on-call and incident management systems, coordinate cloud service deployments.
Seniority
Senior, hands-on IC with mentorship responsibilities
