SRE | Site Reliability Engineering
Core
Ensure reliability, scalability, and evolution of the bank's infrastructure supporting products and operations, combining SRE, Platform Engineering, and AI-driven automation.
Role type
Senior Site Reliability Engineer (SRE)
Builds
High-availability platforms and distributed systems for a digital bank
Domain
Fintech / Cloud Infrastructure / Generative AI
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Service Mesh, distributed systems architecture, observability (OpenTelemetry, Prometheus, Grafana), Infrastructure as Code (Terraform), CI/CD pipelines (ArgoCD, GitHub Actions), Go/Python/Java, SRE principles (SLOs, SLIs, Error Budgets)
Preferred skills
LLMs and AI agents, AIOps, MLOps, FinOps, open source contributions
Technologies
Kubernetes, Service Mesh, OpenTelemetry, Prometheus, Grafana, AppDynamics, Splunk, Terraform, ArgoCD, GitHub Actions
Responsibilities
Maintain high availability and operational excellence of platforms, lead incident response during critical situations, optimize detection and prevention of incidents using AI agents, influence technical and architectural decisions, translate business challenges into sustainable technical solutions
Seniority
Senior, hands-on IC