Site Reliability Engineer (Hosted Infra) - Platform
Core
Cloud Infrastructure SRE integrating, scaling, and evolving multi-cloud infrastructure across 4 CSPs, 70+ regions, and tens of thousands of hosts to power Elastic Cloud.
Role type
Senior IC Site Reliability Engineer (Hosted Infra)
Builds
Internal tools and services to automate large-scale systems; multi-cloud compute infrastructure.
Domain
Cloud Infrastructure / Multi-cloud / Search AI Platform
Deliverable
production ML models | infrastructure
Required skills
Golang, Linux systems, containerized workloads, automation, Infrastructure as Code (IaC), observability, incident response, mentoring
Preferred skills
Terraform, Puppet, Ansible, Argo CD, Argo Workflows, CUE, Docker, Kubernetes, Ubuntu, Elastic Stack, Prometheus, Influx
Technologies
Golang, Linux, Kubernetes, Docker, Terraform, Puppet, Ansible, Argo CD, Argo Workflows, CUE, Elastic Stack, Prometheus, Influx
Responsibilities
Engineering software to automate large-scale systems; optimizing host reliability and lifecycle across multiple cloud providers; strengthening observability posture; scaling global infrastructure; participating in on-call rotation and incident response; mentoring teammates.
Seniority
Senior, hands-on IC