CareerPlanSign in

Site Reliability Engineer

APAC🌐 Remote💼 Full-time🗓 2026-08-10 → 2026-09-26

Core

Design, build, and operate reliable production infrastructure supporting AI Co-Workers.

Role type

Site Reliability Engineer (Infrastructure & Platform)

Builds

Kubernetes-based platforms for AI workloads, infrastructure as code, and observability systems.

Domain

AI infrastructure, Cloud-native platforms

Deliverable

production ML models | infrastructure

Required skills

Kubernetes (EKS/AKS/GKE), Terraform, Helm, CI/CD, Cloud provider (AWS/Azure/GCP), GitOps (ArgoCD), Python/Go/Java/Bash/PowerShell/Ruby

Preferred skills

CKA/CKAD certification, DevSecOps experience

Technologies

Kubernetes, Terraform, Helm, ArgoCD, AWS, Azure, Google Cloud

Responsibilities

Design and deploy Kubernetes-based platforms for AI workloads; Build and maintain infrastructure as code using Terraform; Implement and maintain Helm-based deployment workflows; Build and improve observability across monitoring, logging, and alerting; Partner with engineers to ensure systems are resilient, scalable, and secure.

Seniority

Mid-Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.