Staff Site Reliability Engineer - Kubernetes
Core
Building and managing Kubernetes platforms on AWS to support cloud-native applications, ensuring high availability, performance, and security for the Workforce Identity Cloud.
Role type
Staff Site Reliability Engineer (Kubernetes Platform)
Builds
Highly available, scalable, and secure Kubernetes-based platforms on AWS
Domain
Cloud Infrastructure / Kubernetes / Identity Security
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes platform creation and management, Terraform, AWS infrastructure, Helm, Karpenter, Istio service mesh, CI/CD pipelines, Python/Bash/Go scripting, monitoring and logging tools
Preferred skills
Docker containerization, cloud security best practices (RBAC, encryption)
Technologies
Kubernetes, AWS (EKS, ECS, S3, VPC, RDS, IAM), Helm, Karpenter, Istio, Terraform, Jenkins, GitLab, CircleCI, Ansible, Spinnaker, Prometheus, Grafana, CloudWatch, ELK Stack
Responsibilities
Design and maintain highly available Kubernetes clusters; Build and optimize AWS cloud infrastructure; Automate deployments and scaling using Helm and Karpenter; Manage Istio service mesh for traffic and security; Respond to incidents and troubleshoot system issues; Implement secure cloud infrastructure with compliance frameworks; Create and maintain operational documentation.
Seniority
Staff, hands-on IC with strategic ownership