CareerPlanSign in

Site Reliability Engineer - Cloud Operations

Gland, VD, ch💼 Full-time🗓 2026-08-28 → 2026-09-25

Core

Migrate and modernize production applications on Kubernetes, operate service mesh platforms, and integrate AI tools for troubleshooting and operational automation.

Role type

Senior Site Reliability Engineer (Cloud Operations)

Builds

Reliable, modernized production platforms and cloud-native application infrastructure

Domain

Financial services / Cloud-native infrastructure / AI observability

Deliverable

production ML models | product features | infrastructure

Required skills

Kubernetes (EKS/OpenShift), Helm, Service Mesh (Istio/Linkerd), GitOps, SRE concepts (SLOs/SLIs), Observability (Prometheus/Grafana/OpenTelemetry), Linux/Networking, Python/Go/Bash, JVM troubleshooting, Infrastructure as Code (Terraform/Ansible)

Preferred skills

Argo CD/Rollouts, eBPF/Cilium, GPU/AI infrastructure management, Java/Spring Boot tuning, Public cloud environments, Kubernetes certifications (CKAD/CKS)

Technologies

Kubernetes, OpenShift, EKS, Helm, Istio, Linkerd, Prometheus, Grafana, Elastic Stack, OpenTelemetry, Terraform, Ansible, Python, Go, Bash, vLLM, NVIDIA MIG

Responsibilities

Migrate and modernize production applications on Kubernetes, Integrate third-party software into production platforms, Design and operate applications on service mesh platform, Implement safe deployment patterns (canary/progressive rollouts), Define and use SLOs/SLIs to drive improvements, Improve observability across metrics, logs and traces, Automate repetitive operational work, Provide Level-3 support and participate in on-call rotation

Seniority

Senior, hands-on IC

Sourced via smartrecruiters · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.