Site Reliability Engineer- SB25
Core
Own reliability, scalability, and day-2 operations of Kubernetes platforms including SUSE Harvester, Longhorn, and Rancher-managed clusters.
Role type
Site Reliability Engineer (Kubernetes Platform)
Builds
Production Kubernetes clusters, VM workloads via KubeVirt, and cloud/on-prem resources via Crossplane.
Domain
Cloud Infrastructure / Kubernetes / HCI
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Kubernetes administration, SUSE Harvester, Longhorn, Antrea CNI, KubeVirt, Rancher, FluxCD, Crossplane, Linux networking, observability, scripting (Bash/Python), HA design, capacity planning.
Preferred skills
GitOps patterns, Kustomize/Helm, TLS/cert management, disaster recovery.
Technologies
SUSE Harvester, Longhorn, Rancher, FluxCD, Crossplane, Antrea, KubeVirt, CAPI, Helm, Kustomize.
Responsibilities
Operate and harden SUSE Harvester environments, manage multi-cluster Kubernetes lifecycles, own CNI operations with Antrea, run KubeVirt for VM workloads, implement GitOps with FluxCD, provision resources with Crossplane, build and maintain SLOs/SLIs, reduce toil through automation, participate in on-call rotations and incident reviews.
Seniority
Mid-level, hands-on IC
