Senior Site Reliability Engineer, Kubernetes w/ active TS/SCI
Core
Design, deploy, and monitor large-scale cloud production infrastructure to ensure peak performance and reliability for sensitive national security missions.
Role type
Senior Site Reliability Engineer (Kubernetes)
Builds
Okta's production infrastructure and automation tools
Domain
Cloud Infrastructure / National Security
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Linux systems administration, Go/Python/Bash/Ruby, AWS services, Infrastructure as Code (Terraform/CloudFormation), Docker, incident management, automation scripting
Preferred skills
Multi-cloud environments, Helm chart debugging
Technologies
Kubernetes, AWS, Terraform, CloudFormation, Docker, Helm, Go, Python, Bash, Ruby
Responsibilities
Design and deploy production infrastructure; respond to production incidents and implement preventive solutions; develop automation scripts to eliminate manual toil; support high-availability environments via on-call rotation
Seniority
Senior, hands-on IC