Senior Site Reliability Engineer - Managed Kubernetes
Core
Operate and maintain bare-metal Kubernetes clusters for AI/ML workloads, handling scaling, incident response, and customer support.
Role type
Senior Site Reliability Engineer (Managed Kubernetes)
Builds
Scalable control plane services, operators, custom controllers, and automation for cluster lifecycle management.
Domain
Cloud Infrastructure / AI/ML / Kubernetes
Deliverable
production ML models | infrastructure
Required skills
Kubernetes cluster operations, Go, Python, GitOps (ArgoCD), Helm, Kubernetes operators, observability (Prometheus, Grafana), CI/CD pipelines, cluster provisioning (kubeadm, Cluster API)
Preferred skills
CRDs, CSI, CNI, Kubernetes Operator Coding, HPC clusters, AI/ML workloads, large-scale GPU clusters, hybrid/multi-cloud environments, CNCF contributions
Responsibilities
Operate and maintain bare-metal Kubernetes clusters scaling to thousands of nodes; Handle cluster degradation, recovery, resizing, and incident response; Participate in on-call rotation for critical incidents; Assist customers with Kubernetes questions and workload integration; Design and maintain scalable control plane services and operators; Develop automation for cluster lifecycle management; Define and implement SLOs and SLIs for platform reliability.
Seniority
Senior, hands-on IC