Senior Site Reliability Engineer - Platform Reliability (Resilience)
Core
Designing, building, scaling, and maturing the multi-cloud platform for hosting internal and external services like Elastic Cloud Hosted and Serverless to ensure global infrastructure reliability.
Role type
Senior Site Reliability Engineer (Platform Reliability)
Builds
Multi-cloud platform infrastructure, automation tooling, and software supporting product deployment across Elastic Cloud Hosted and Serverless.
Domain
Cloud Infrastructure / SaaS / Search & Security
Deliverable
production ML models | infrastructure
Required skills
Software engineering, public cloud platforms, managed Kubernetes services, Infrastructure-as-Code (IaC), Golang or other programming languages, containerized services (Docker), Linux system administration on distributed systems, alerting and incident management processes.
Preferred skills
Experience operating SaaS products in public cloud, Kubernetes-at-scale across multiple providers, leading alerting/metrics systems (Elastic Stack, Prometheus, Influx), coaching and mentoring team members.
Responsibilities
Lead technical initiatives to automate system engineering for global infrastructure reliability, develop and maintain software/tooling for infrastructure scaling, respond to and prevent major customer impact incidents, champion collaboration and operational excellence.
Seniority
Senior, hands-on IC
