Site Reliability Engineer
Core
Bridge software development and IT operations to build scalable, reliable, and automated systems for critical business solutions.
Role type
Site Reliability Engineer (SRE)
Builds
Scalable, reliable, and automated infrastructure and operational systems
Domain
Financial risk services (Anti-Money Laundering, Fraud, Credit Risk) on cloud platforms
Deliverable
production ML models | infrastructure
Required skills
Python, Go, Java, Bash, Linux/Unix, TCP/IP networking, AWS/GCP/Azure, Prometheus, Grafana, CI/CD, Ansible, Terraform
Preferred skills
None stated
Technologies
Python, Go, Java, Bash, Linux, AWS, GCP, Azure, Prometheus, Grafana, Ansible, Terraform
Responsibilities
Develop scripts to automate provisioning, configuration, and incident response; Monitor system performance and establish alerting; Reduce MTTR through troubleshooting and operational improvements; Scale infrastructure and support capacity planning; Analyze logs, traces, and metrics to optimize performance; Manage error budgets and define/maintain SLIs/SLOs
Seniority
Mid-level IC