Site Reliability Engineer
Core
Design, build, and operate reliable production infrastructure supporting AI Co-Workers.
Role type
Site Reliability Engineer (Infrastructure & Platform)
Builds
Kubernetes-based platforms for AI workloads
Domain
AI infrastructure / Cloud native
Deliverable
production ML models
Required skills
Kubernetes, Terraform, Helm, Python, Go, Java, Bash, PowerShell, Ruby, CI/CD, observability, incident response, root cause analysis, automation
Preferred skills
Cloud provider experience (AWS/Azure/GCP), GitOps/ArgoCD, DevSecOps
Technologies
Kubernetes, Terraform, Helm, ArgoCD, AWS, Azure, Google Cloud
Responsibilities
Design and operate reliable production infrastructure; Own Kubernetes-based platforms; Build and maintain infrastructure as code; Implement Helm-based deployment workflows; Define and improve system reliability using SLIs/SLOs/SLAs; Participate in on-call rotation and incident response; Reduce operational toil through automation; Build and improve observability; Partner with engineers on resilience and security.
Seniority
Mid-Senior, hands-on IC
