Site Reliability Engineer
Core
Maintain reliability, scalability, and performance of production systems through automation, monitoring, and incident management.
Role type
Site Reliability Engineer (SRE)
Builds
Automated deployment, monitoring, and alerting solutions; scalable infrastructure via IaC
Domain
Financial services / Cloud infrastructure
Deliverable
production ML models | infrastructure
Required skills
Ansible, Terraform, Jenkins, Grafana, AWS/Azure/GCP, Docker, Kubernetes, CI/CD pipelines, Python/Bash/Go
Preferred skills
Prometheus, networking and security best practices, database administration and optimization
Technologies
Ansible, Terraform, Jenkins, Grafana, Prometheus, AWS, Azure, Google Cloud, Docker, Kubernetes
Responsibilities
Design and manage automated deployment, monitoring, and alerting solutions; Build and support scalable infrastructure through Infrastructure as Code (IaC) tools; Use Grafana and other monitoring platforms to track system reliability and performance; Diagnose and resolve production issues quickly to minimize downtime; Create and maintain best practices and guidelines for SRE processes; Enhance observability by improving logging, monitoring, and alert systems; Participate in on-call rotations to ensure round-the-clock support for critical systems; Lead post-incident reviews and put preventative measures in place; Mentor and educate team members on SRE methodologies and technologies
Seniority
Mid-to-Senior, hands-on IC