Senior Site Reliability Engineer (Capacity) - Platform Infrastructure
Core
Manage and optimize compute resources for Elastic Cloud Hosted and Serverless workloads to ensure seamless scaling across 60+ regions.
Role type
Principal Platform Engineer (Capacity)
Builds
Elastic Cloud Hosted and Serverless infrastructure
Domain
Cloud Infrastructure / Search AI Platform
Deliverable
production ML models | infrastructure
Required skills
cloud infrastructure capacity management, performance monitoring and optimization, cloud scaling strategies, incident investigation and troubleshooting, compute auto-scaling, capacity reservations, software and platform engineering
Preferred skills
experience with the three major cloud service providers (AWS, Azure, GCP)
Technologies
Elastic Cloud, autoscaling frameworks, cloud CSPs (AWS, Azure, GCP)
Responsibilities
Assess current and future capacity requirements based on workload demands, develop and maintain accurate capacity models, implement strategies for optimizing resource usage, analyze capacity metrics and trends, operate an autoscaling framework across 60+ regions
Seniority
Principal, hands-on IC