Site Reliability Engineer
Core
Design, build, and operate reliable production infrastructure supporting AI Co-Workers.
Role type
Site Reliability Engineer (Infrastructure & Platform)
Builds
Kubernetes-based platforms for AI workloads, infrastructure as code, and observability systems.
Domain
AI infrastructure, Cloud-native platforms
Deliverable
production ML models | infrastructure
Required skills
Kubernetes (EKS/AKS/GKE), Terraform, Helm, CI/CD, Cloud provider (AWS/Azure/GCP), GitOps (ArgoCD), Python/Go/Java/Bash/PowerShell/Ruby
Preferred skills
CKA/CKAD certification, DevSecOps experience
Technologies
Kubernetes, Terraform, Helm, ArgoCD, AWS, Azure, Google Cloud
Responsibilities
Design and deploy Kubernetes-based platforms for AI workloads; Build and maintain infrastructure as code using Terraform; Implement and maintain Helm-based deployment workflows; Build and improve observability across monitoring, logging, and alerting; Partner with engineers to ensure systems are resilient, scalable, and secure.
Seniority
Mid-Senior, hands-on IC
