Senior Site Reliability Engineer
Core
Design, build, and maintain highly reliable and resilient infrastructure platforms for large-scale on-premises environments, driving automation and operational efficiency.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Highly reliable and resilient infrastructure platforms
Domain
Financial markets infrastructure and data provider
Deliverable
production ML models | infrastructure
Required skills
Python (Advanced), Ansible, Terraform, CI/CD platforms, Infrastructure as Code (IaC), Observability, Monitoring, Alerting, Troubleshooting, SRE concepts (reliability, availability, capacity planning), ITIL principles (Incident, Change, Problem Management), Root Cause Analysis (RCA)
Preferred skills
OpenTelemetry (OTEL), Grafana, Cribl, ClickHouse, enterprise-scale on-prem or hybrid environment support
Technologies
Python, Ansible, Terraform, OpenTelemetry, Grafana, Cribl, ClickHouse
Responsibilities
Design, build, and maintain highly reliable and resilient infrastructure platforms; Drive automation initiatives using Python, Ansible, and Terraform; Develop and support CI/CD pipelines and Infrastructure as Code (IaC) practices; Implement and enhance observability solutions, monitoring, alerting, and troubleshooting frameworks; Lead reliability improvements, incident response, RCA activities, and service restoration efforts; Apply ITIL best practices for Incident, Change, and Problem Management to ensure operational governance and service stability.