Site Reliablity Engineer III
Core
Build, deploy, scale, and operate highly reliable distributed systems across multiple geographic regions to ensure high availability and performance.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Kubernetes-based infrastructure for large-scale applications and cloud-native ecosystems
Domain
Cybersecurity, Cloud Infrastructure, Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Terraform, AWS (EKS, VPC, S3, ECR, IAM), Helm charts, Bash scripting, CI/CD pipelines, observability stacks (Grafana, Prometheus, Loki), debugging and root cause analysis
Preferred skills
Golang, Python, multi-region deployments, global infrastructure management
Technologies
Kubernetes, Terraform, AWS, Helm, Grafana, Prometheus, Loki, Bash
Responsibilities
Deploy and manage distributed platforms across multiple regions; Design and maintain Kubernetes-based infrastructure; Monitor system health and resolve issues; Implement automation and best practices for reliability; Collaborate on CI/CD and production readiness; Troubleshoot production issues and write run books
Seniority
Senior, hands-on IC