CareerPlanSign in

Senior Site Reliability Engineer - Managed Kubernetes

San Francisco Office (Fremont St)💼 Full-time🗓 2026-08-14 → 2026-09-25

Core

Operate and maintain bare-metal Kubernetes clusters for AI/ML workloads, handling scaling, incident response, and customer support.

Role type

Senior Site Reliability Engineer (Managed Kubernetes)

Builds

Scalable control plane services, operators, custom controllers, and automation for cluster lifecycle management.

Domain

Cloud Infrastructure / AI/ML / Kubernetes

Deliverable

production ML models | infrastructure

Required skills

Kubernetes cluster operations, Go, Python, GitOps (ArgoCD), Helm, Kubernetes operators, observability (Prometheus, Grafana), CI/CD pipelines, cluster provisioning (kubeadm, Cluster API)

Preferred skills

CRDs, CSI, CNI, Kubernetes Operator Coding, HPC clusters, AI/ML workloads, large-scale GPU clusters, hybrid/multi-cloud environments, CNCF contributions

Responsibilities

Operate and maintain bare-metal Kubernetes clusters scaling to thousands of nodes; Handle cluster degradation, recovery, resizing, and incident response; Participate in on-call rotation for critical incidents; Assist customers with Kubernetes questions and workload integration; Design and maintain scalable control plane services and operators; Develop automation for cluster lifecycle management; Define and implement SLOs and SLIs for platform reliability.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.