Site Reliability Engineer
Core
Design, build, and operate reliable production infrastructure supporting AI Co-Workers that automate real operational work for global enterprises.
Role type
Site Reliability Engineer (SRE)
Builds
Kubernetes-based platforms for deploying and running AI workloads
Domain
AI infrastructure / Cloud Operations
Deliverable
production ML models
Required skills
Kubernetes, Terraform, Helm, Python/Go/Java/Bash/PowerShell/Ruby, Cloud provider (AWS/Azure/GCP), ArgoCD/GitOps, CI/CD, DevSecOps
Preferred skills
CKA, CKAD, cloud certifications, DevOps/DevSecOps certifications
Technologies
Kubernetes, EKS, AKS, GKE, Terraform, Helm, ArgoCD, AWS, Azure, Google Cloud
Responsibilities
Design and operate reliable production infrastructure; Own Kubernetes-based platforms; Build and maintain infrastructure as code; Implement Helm-based deployment workflows; Define and improve system reliability using SLIs/SLOs/SLAs; Participate in on-call rotation and incident response; Reduce operational toil through automation; Build and improve observability; Partner with engineers to ensure system resilience and security.
Seniority
Mid-Senior, hands-on IC