Site Reliability Engineer - Cloud Operations
Core
Migrate and modernize production applications on Kubernetes, operate service mesh platforms, and integrate AI tools for troubleshooting and operational automation.
Role type
Senior Site Reliability Engineer (Cloud Operations)
Builds
Reliable, modernized production platforms and cloud-native application infrastructure
Domain
Financial services / Cloud-native infrastructure / AI observability
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes (EKS/OpenShift), Helm, Service Mesh (Istio/Linkerd), GitOps, SRE concepts (SLOs/SLIs), Observability (Prometheus/Grafana/OpenTelemetry), Linux/Networking, Python/Go/Bash, JVM troubleshooting, Infrastructure as Code (Terraform/Ansible)
Preferred skills
Argo CD/Rollouts, eBPF/Cilium, GPU/AI infrastructure management, Java/Spring Boot tuning, Public cloud environments, Kubernetes certifications (CKAD/CKS)
Technologies
Kubernetes, OpenShift, EKS, Helm, Istio, Linkerd, Prometheus, Grafana, Elastic Stack, OpenTelemetry, Terraform, Ansible, Python, Go, Bash, vLLM, NVIDIA MIG
Responsibilities
Migrate and modernize production applications on Kubernetes, Integrate third-party software into production platforms, Design and operate applications on service mesh platform, Implement safe deployment patterns (canary/progressive rollouts), Define and use SLOs/SLIs to drive improvements, Improve observability across metrics, logs and traces, Automate repetitive operational work, Provide Level-3 support and participate in on-call rotation
Seniority
Senior, hands-on IC
