Site Reliability Engineer - Multicloud Platform
Core
Ensure reliability and availability of a multicloud platform, reduce operational toil, and scale sustainably for Workday's cloud-native services.
Role type
Senior Site Reliability Engineer (Multicloud Platform)
Builds
Cloud-native technology platform on AWS, GCP, and Azure supporting Workday products
Domain
Cloud Native / Multicloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Linux, distributed systems, public cloud (AWS/GCP/Azure), GoLang/Python/Ruby, CI/CD, observability (SLIs/SLOs), automation
Preferred skills
Cloud Native Conference experience, independent problem solving, fast-paced environment adaptability
Technologies
Kubernetes, Istio, OPA, GoLang, Prometheus, Grafana, AWS, GCP, Azure
Responsibilities
Develop and launch effective SLIs to ensure SLOs are achieved; Partner with platform service teams to craft and implement SRE standards; Define benchmarks and automation to qualify services for production; Ensure safe change and reliability of customer environments; Operate and maintain the multicloud platform
Seniority
Senior, hands-on IC