Software Engineer - Site Reliability
Core
Building reliable, scalable systems and processes for the Workday Data Platform, focusing on infrastructure, automation, and CI/CD pipelines to support backend, frontend, and platform engineering teams.
Role type
Senior Site Reliability Engineer (Infrastructure & Automation)
Builds
Immutable services, functions, and automated deployment pipelines for the Prism data platform.
Domain
Cloud Infrastructure (AWS/GCP) and Big Data Processing
Deliverable
infrastructure
Required skills
Distributed systems design, Unix/Linux kernel to shell, Python, GoLang, Java, Docker, Kubernetes, Terraform, Ansible, CI/CD pipeline creation, JVM debugging and tuning
Preferred skills
Spark, YARN, Hadoop, Trino, Iceberg, Polaris, Prometheus, Grafana, Serverless frameworks (AWS Lambda, API Gateway)
Technologies
AWS, GCP, Spark, YARN, Hadoop, Kubernetes, Docker, Terraform, Ansible, Jenkins, TeamCity, Bamboo, Artifactory, Prometheus, Grafana, Python, GoLang, Java
Responsibilities
Design, analyze, and troubleshoot large-scale distributed systems; automate software development lifecycles and deployment steps; build tooling and infrastructure in the cloud; create meaningful metrics and alerts for system health; debug and tune JVM applications.
Seniority
Senior, hands-on IC