Senior Site Reliability Engineer (Observability & Analytics) – Platform Infra
Core
Building and maintaining the observability infrastructure (logs, metrics, traces) and analytics pipelines for Elastic Cloud, enabling Cloud engineers to monitor production status and the business to analyze platform usage.
Role type
Senior Site Reliability Engineer (Observability & Analytics)
Builds
200+ hosted deployments across cloud regions, ingesting logs/metrics/traces for Elastic Cloud, and SLA/SLO monitoring for ESS and Serverless.
Domain
Cloud Infrastructure / Observability / SaaS Platform
Deliverable
production ML models | infrastructure
Required skills
Terraform, Python, Go, Linux systems, containerized workloads, incident response, RCA writing, code review, mentoring, security-conscious infrastructure design
Preferred skills
Elastic Stack (Elasticsearch, Logstash, Beats, Kibana), GitOps (ArgoCD, Helm), policy-as-code (Kyverno), Kubernetes, secrets management (Vault), access-control (Teleport), configuration management (Puppet, Ansible), FedRAMP/GovCloud experience
Responsibilities
Own end-to-end delivery of moderate-to-high complexity projects; operate and harden shared Elastic Cloud infrastructure as IaC; carry 24/7 on-call rotation for incident resolution; review code and designs; mentor less experienced engineers; improve runbooks and operational processes
Seniority
Senior, hands-on IC