Site Reliability Engineer
Core
Ensure reliability, availability, performance, and scalability of business-critical applications and platforms through production support, cloud operations, and automation.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Production support, automation solutions, observability pipelines, and cloud infrastructure
Domain
Financial markets infrastructure, cloud operations, DevOps
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Python, PowerShell, Shell scripting, Linux/Unix, Windows, Azure, AWS, Terraform, SQL, incident management, RCA, SLOs, monitoring, log analysis
Preferred skills
Kubernetes, containers, microservices, CI/CD pipelines, Kafka, APIs, middleware, cloud-native databases, Identity & Access Management, ITIL certifications
Technologies
Python, PowerShell, Shell, Terraform, Datadog, BigPanda, Splunk, Azure, AWS, Kafka
Responsibilities
Own and improve service reliability through SLOs and proactive problem management; Provide production support and resolve complex incidents within SLA; Conduct incident management, RCA, and postmortems; Develop automation and self-healing solutions; Implement and support Infrastructure as Code and CI/CD practices; Design and maintain observability solutions; Support cloud platforms and migration initiatives; Execute Operational Acceptance Testing; Participate in on-call rotations and major incident response.
Seniority
Senior, hands-on IC