Senior Site Reliability Engineer (US Federal)
Core
Building a user-friendly, scalable, and reliable tools framework for the health, performance, and reliability of the Workday application in the public cloud, specifically supporting U.S. Federal Government contracts.
Role type
Senior Site Reliability Engineer (Public Cloud Operations)
Builds
Scalable tools framework, automated cluster buildouts, validation, and patch deployment pipelines
Domain
Public Cloud Infrastructure, U.S. Federal Government Contracts
Deliverable
production ML models | infrastructure
Required skills
Python, Ruby, GoLang, Java, Terraform, Ansible, Chef, Docker, Kubernetes, Serverless (Lambda), AWS, Google Cloud Platform, Linux, Shell Scripting, SQL, MySQL, CI tools (Jenkins, TeamCity, Bamboo, Artifactory)
Preferred skills
DoD 8570/8140 compliance (IAT Level II), CompTIA CySA+, GICSP, CASP+, TS/SCI security clearance
Technologies
Docker, Kubernetes, Serverless, AWS, Google Cloud Platform, Terraform, Ansible, Chef, Jenkins, TeamCity, Bamboo, Artifactory
Responsibilities
Automate cluster buildouts, validation, patch deployment, and operational tasks; Design and analyze large-scale distributed systems; Mentor and guide fellow team members; Collaborate with other Environments teams to ensure seamless operations
Seniority
Senior, hands-on IC with mentorship responsibilities