Site Reliability Engineer
Core
Ensuring reliability, availability, performance, and scalability of fintech infrastructure and applications through automation and cloud management.
Role type
Site Reliability Engineer
Builds
Resilient, secure, and efficient platforms supporting continuous delivery for SME financial services.
Domain
Fintech / Cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes administration, Infrastructure as Code (Terraform, Ansible), Cloud provider expertise (AWS, GCP, Azure, OCI), CI/CD pipeline design (GitOps), Monitoring and observability (Prometheus, Loki, Jaeger, Sentry), Database management (PostgreSQL, MongoDB), System architecture and troubleshooting, Networking (VPNs, Service Mesh), Root cause analysis, Automation tooling.
Preferred skills
Compliance and regulations knowledge, Service Mesh experience, VPN configuration.
Responsibilities
Monitor cloud environment availability and system health, Build software and systems to manage platform infrastructure, Improve reliability and time-to-market of software solutions, Measure and optimize system performance, Provide primary operational support for distributed applications, Analyze metrics for performance tuning and fault finding, Partner with development teams on testing and release procedures, Participate in system design consulting and capacity planning, Create sustainable systems through automation, Deploy updates and fixes, Build tools to reduce errors, Perform root cause analysis for production errors, Investigate and resolve technical issues, Design system troubleshooting procedures.
Seniority
Mid-level, hands-on IC