Site Reliability Engineer
Core
Own production readiness, resilience, and operational health of cutting-edge AI solutions, bridging application quality, MLOps, and cloud resilience for banking-scale traffic.
Role type
Senior Site Reliability Engineer (AI/MLOps focus)
Builds
Automated test suites for LLM/RAG outputs, chaos-resilient cloud environments, observability frameworks, and security/privacy automation pipelines.
Domain
Financial Services / Banking / AI & Data Consultancy
Deliverable
production ML models | infrastructure
Required skills
Python, Bash scripting, PyTest/Selenium/Robot Framework, Azure/AWS cloud networking, containerized environments, serverless reliability, CI/CD (GitHub Actions/Azure DevOps), Terraform, K6/JMeter, Azure Monitor/CloudWatch/Grafana/Prometheus, prompt engineering, RAG architectures, model evaluation metrics (ROUGE/BLEU, LangSmith/Ragas).
Preferred skills
Kubernetes cluster management (AKS/EKS), service meshes, container security scanners, UAE Central Bank regulatory guidelines, NESA security compliance.
Technologies
Python, Bash, PyTest, Selenium, Robot Framework, Azure, AWS, GitHub Actions, Azure DevOps, Terraform, K6, JMeter, Azure Monitor, CloudWatch, Grafana, Prometheus, DeepEval, Ragas, LangSmith, Kubernetes, AKS, EKS.
Responsibilities
Implement load testing, fault-injection techniques, and self-healing automation for web/mobile backends and AI microservices; Build Python-driven, CI/CD-integrated regression test suites validating application logic, data pipelines, and IaC states; Define and track critical SLIs/SLOs and set up real-time telemetry, dashboards, and proactive alerting; Integrate specialized AI evaluation harnesses into deployment pipelines to benchmark model accuracy, hallucinations, and retrieval performance; Automate security scans and privacy checks to satisfy strict banking data residency standards while monitoring compute efficiency.
Seniority
Senior, hands-on IC