Site Reliability Engineer (SRE)
Core
Ensure the reliability, scalability, and performance of a global fintech platform serving businesses and AI agents through cloud infrastructure and distributed systems.
Role type
Senior IC Site Reliability Engineer (SRE)
Builds
Cloud-based infrastructure, distributed systems, and automation tooling for a global financial ecosystem.
Domain
Fintech / Cloud Infrastructure / Distributed Systems
Deliverable
production ML models | product features | infrastructure
Required skills
Linux systems engineering, Kubernetes, AWS/GCP, Terraform, Ansible, Python/Golang/bash scripting, distributed systems design, observability tooling (Prometheus/Grafana/PagerDuty), CI/CD pipeline management
Preferred skills
High-volume web services operation, SLO/SLI management, AI-driven mindset, AI tool usage
Technologies
Linux, Kubernetes, GitLab, Terraform, Ansible, AWS, Google Cloud Platform, Prometheus, Grafana, Zabbix, Splunk, PagerDuty, PostgreSQL, MongoDB, RabbitMQ, Apache, Nginx
Responsibilities
Ensure platform availability and resilience; own incident management and on-call rotation; design and manage Linux-based system architecture; build and support Kubernetes workloads; implement Infrastructure as Code (IaC); manage CI/CD pipelines; implement monitoring, logging, and alerting; produce operational documentation and standards.
Seniority
Senior, hands-on IC