Staff SRE for K8s Platform Team (AWS, Kubernetes, Platform Creation, Helm, Karpenter, Istio)
Core
Architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS to support cloud-native applications and services.
Role type
Staff Site Reliability Engineer (Kubernetes Platform)
Builds
Highly available, scalable, and fault-tolerant Kubernetes clusters and AWS infrastructure
Domain
Cloud-native infrastructure, Kubernetes, AWS
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes platform creation and management, Helm chart management, Karpenter implementation, Istio service mesh management, AWS infrastructure management, Infrastructure-as-code (Terraform, Chef, Ansible), Serverless computing (AWS Lambda, API Gateway), CI/CD pipeline automation, Python or Go scripting, Monitoring and logging (Prometheus, Grafana, CloudWatch, ELK)
Preferred skills
Multi-region cloud environment experience, Cloud security best practices (RBAC, encryption), Docker containerization
Technologies
Kubernetes, AWS (EKS, ECS, S3, VPC, RDS, IAM), Helm, Karpenter, Istio, Terraform, Chef, Ansible, AWS Lambda, API Gateway, Jenkins, GitLab, CircleCI, Spinnaker, Prometheus, Grafana, CloudWatch, ELK Stack
Responsibilities
Design and maintain highly available Kubernetes clusters; Build and optimize AWS cloud infrastructure; Automate application deployment using Helm; Implement dynamic cluster scaling with Karpenter; Configure service mesh with Istio; Manage CI/CD pipelines; Respond to incidents and troubleshoot system issues; Design secure cloud infrastructure; Create operational documentation
Seniority
Staff, hands-on IC