Site Reliability Engineer - eFX/ Crypto
Core
Ensure reliability, availability, scalability, and performance of production eFX/Crypto platforms using Kubernetes and automation.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Production eFX/Crypto platforms on on-premise and private cloud infrastructure
Domain
Financial services (eFX, Crypto) + Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes (production), Linux administration, Docker, Infrastructure as Code (Terraform, Ansible), CI/CD pipelines, Observability (Prometheus, Grafana, Elastic/Kibana), Python, Shell scripting, Networking fundamentals (TCP/IP, DNS, HTTP, Load Balancing), SRE principles (SLIs, SLOs, SLAs)
Preferred skills
Istio, Helm, ArgoCD, Kafka, RabbitMQ, MT4/MT5, Windows Server, High-availability systems, eFX/trading/crypto/banking domain experience
Technologies
Kubernetes, Docker, Terraform, Ansible, Prometheus, Grafana, Elastic/Kibana, Istio, Python, Shell, Spring Boot, Apache Tomcat
Responsibilities
Monitor production systems and respond to incidents, Operate and troubleshoot Kubernetes clusters and containerized applications, Automate IT and operational processes, Implement and maintain CI/CD pipelines, Perform root-cause analysis and implement corrective actions, Conduct stress, resilience, disaster recovery, and BCP testing
Seniority
Senior, hands-on IC