CareerPlanSign in

Site Reliability Engineer

Austin💼 Full-time🗓 2026-07-07 → 2026-09-26

Core

Design, build, and operate reliable production infrastructure supporting AI Co-Workers.

Role type

Site Reliability Engineer (Infrastructure & Platform)

Builds

Kubernetes-based platforms for AI workloads

Domain

AI infrastructure / Cloud native

Deliverable

production ML models

Required skills

Kubernetes, Terraform, Helm, Python, Go, Java, Bash, PowerShell, Ruby, CI/CD, observability, incident response, root cause analysis, automation

Preferred skills

Cloud provider experience (AWS/Azure/GCP), GitOps/ArgoCD, DevSecOps

Technologies

Kubernetes, Terraform, Helm, ArgoCD, AWS, Azure, Google Cloud

Responsibilities

Design and operate reliable production infrastructure; Own Kubernetes-based platforms; Build and maintain infrastructure as code; Implement Helm-based deployment workflows; Define and improve system reliability using SLIs/SLOs/SLAs; Participate in on-call rotation and incident response; Reduce operational toil through automation; Build and improve observability; Partner with engineers on resilience and security.

Seniority

Mid-Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.