Site Reliability Engineer
Core
Build end-to-end observability solutions using Azure Monitor, ADX (KQL), and Grafana to support Payments and Order to Cash reliability workflows.
Role type
Site Reliability Engineer (SRE) with focus on AI-driven engineering and observability
Builds
Scalable observability stacks, automated diagnostics, and self-healing mechanisms for mission-critical payment systems
Domain
Financial services / Payments / Cloud Infrastructure (Azure)
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Azure platform services, ADX (KQL), Grafana, Python, PowerShell, C#, Infrastructure as Code (IaC), CI/CD pipelines, Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, AI tools (M365, GitHub Copilot), automated diagnostics, capacity planning, performance optimization
Preferred skills
Vendor-driven operation team experience, blameless Root Cause Analysis (RCA), load testing, smart alerting strategies
Technologies
Azure Monitor, ADX, KQL, Grafana, Python, PowerShell, C#, M365, GitHub Copilot
Responsibilities
Define and monitor SLIs, SLOs, and Error Budgets for payment workflows; Architect and maintain scalable observability stacks; Reduce incidents and response times via automation; Design AI-assisted solutions and automated diagnostics; Monitor system health for capacity planning and performance tuning
Seniority
Mid-to-Senior, hands-on IC