Senior Site Reliability Engineer - Platform Reliability (Resilience)
Core
Designing, building, scaling, and maturing the multi-cloud platform for hosting internal and external services like Elastic Cloud Hosted and Serverless to ensure global infrastructure reliability.
Role type
Senior Site Reliability Engineer (Platform Reliability)
Builds
Multi-cloud platform infrastructure, automation tooling, and software supporting product deployment across public clouds.
Domain
Cloud Infrastructure, SaaS, Search AI, Kubernetes
Deliverable
production ML models | infrastructure
Required skills
Software engineering, public cloud platforms, managed Kubernetes services, Infrastructure-as-Code (IaC), Linux system administration, distributed systems, alerting and incident management, mentoring/coaching
Preferred skills
Golang programming, containerized services (Docker), Elastic Stack, cross-cloud Kubernetes operations, self-organizing team environments
Technologies
Kubernetes, Terraform, Crossplane, Elastic Stack, Prometheus, Graphite, Influx, Docker, Linux
Responsibilities
Lead technical initiatives to automate system engineering for global infrastructure reliability; develop and maintain software, tooling, and automations to scale platform infrastructure; respond to and prevent repeated customer impact during major incidents; champion an environment focused on collaboration and operational excellence.
Seniority
Senior, hands-on IC with mentorship responsibilities
