Site Reliability Engineer
Core
Support improvements to availability, performance, efficiency, change management, monitoring, security, incident response, and capacity planning for cloud-native products and services.
Role type
Site Reliability Engineer (SRE)
Builds
Cloud-native platforms on AWS using Kubernetes (EKS)
Domain
Cloud Infrastructure / DevOps
Deliverable
production ML models | product features | infrastructure
Required skills
AWS, Kubernetes (EKS), Terraform, GitOps, CI/CD pipelines, monitoring and observability tools (Grafana, Prometheus), incident response, root cause analysis, automation
Preferred skills
GitLab, Argo CD, networking fundamentals, security compliance awareness
Technologies
AWS, Kubernetes, EKS, Terraform, GitLab, Argo CD, Grafana, Prometheus, Git
Responsibilities
Participate in 24/7 on-call rotations and production support; implement infrastructure changes using Terraform and GitOps; support CI/CD pipelines; assist in incident management and troubleshooting; improve system reliability through automation; follow SRE practices including runbooks and post-incident reviews.
Seniority
Mid-level, hands-on IC