Site Reliability Engineer (4024)
Core
Build and operate the reliability, observability, and operational excellence infrastructure for managed fraud detection platforms serving banking and fintech customers.
Role type
Senior Site Reliability Engineer (SRE)
Builds
High availability SLAs for real-time fraud detection at scale
Domain
Fintech / Fraud Detection / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Cloud platforms (AWS/Azure/GCP), Container orchestration (Kubernetes, Docker, Helm), Infrastructure as Code (Terraform), Observability (Prometheus, Grafana, ELK), CI/CD (Jenkins, GitHub Actions), Security compliance (PCI DSS, ISO 27001, SOC 1/2/3), Database reliability (SQL Server, PostgreSQL, Oracle), Streaming architectures (Kafka), Scripting (Python, Bash, Go)
Preferred skills
Chaos engineering, Secrets management (HashiCorp Vault), Zero-trust identity, Performance testing
Technologies
Terraform, Helm, Kubernetes, AWS, Azure, Jenkins, Prometheus, Grafana, ELK, Kafka, Python, Bash, Go
Responsibilities
Design and operate SRE practices including on-call processes and incident response playbooks; Build and maintain observability infrastructure with centralized logging, metrics, and distributed tracing; Define and track SLOs and error budgets for real-time transaction pipelines; Manage cloud infrastructure provisioning and configuration using IaC; Implement and maintain CI/CD pipelines; Drive platform resilience improvements including auto-scaling and disaster recovery; Manage secrets, identity/access controls, and vulnerability management; Contribute to Architecture Review Committee with operational perspectives.
Seniority
Senior, hands-on IC