Staff Site Reliability Engineer - Cloud Efficiency
Core
Own the cloud efficiency roadmap end-to-end for a platform running thousands of services across multiple Kubernetes clusters and regions.
Role type
Staff Site Reliability Engineer (Cloud Efficiency)
Builds
Cloud efficiency improvements, cost optimization strategies, and Kubernetes resource management at scale.
Domain
Cloud Infrastructure / FinOps / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Large-scale distributed systems operations, cloud cost optimization, architectural judgement, infrastructure as code, AWS proficiency, data store cost/performance trade-offs
Preferred skills
Multi-cloud experience (GCP, Azure), Go/Python automation, CI/CD pipeline policy embedding
Technologies
Kubernetes, AWS, PostgreSQL, Redis, Kafka, DynamoDB, OpenSearch
Responsibilities
Identify high-leverage cloud efficiency opportunities, partner with FinOps to implement changes, surface architectural mismatches causing inefficiencies, own commitment and purchasing strategy, improve guardrails and automation for cost attribution, drive Kubernetes efficiency via bin-packing and autoscaling
Seniority
Staff, hands-on IC with strategic autonomy