Senior Site Reliability Engineer (Capacity) - Platform Infrastructure
Core
Principal Platform Engineer managing compute resources and capacity for Elastic Cloud Hosted and Serverless workloads to ensure seamless scaling.
Role type
Principal Platform Engineer (Capacity)
Builds
Elastic Cloud Hosted and Serverless workloads across 60+ regions
Domain
Cloud Infrastructure / Search AI Platform
Deliverable
production ML models | infrastructure
Required skills
cloud infrastructure capacity management, performance monitoring and optimization, cloud scaling strategies, incident investigation and troubleshooting, compute auto-scaling, capacity reservations, software and platform engineering
Preferred skills
experience with the three major cloud service providers
Technologies
Elastic Cloud, major CSPs (AWS, Azure, GCP)
Responsibilities
Assess current and future capacity requirements based on workload demands, develop and maintain accurate capacity models, implement strategies for optimizing resource usage, analyze capacity metrics and trends, operate an autoscaling framework, collaborate with development teams on scaling best practices
Seniority
Principal, hands-on IC