Senior Site Reliability & Observability Engineer (SRE)
Core
Ensure reliability, stability, and availability of mission-critical services through SRE best practices, observability, and incident management.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Resilient, observable, and automated distributed systems supporting critical business operations.
Domain
Financial services / Distributed systems / Observability
Required skills
Site Reliability Engineering (SRE), distributed systems, observability, incident management, root cause analysis, SLO/SLI management, infrastructure automation, troubleshooting
Preferred skills
Platform engineering, blameless post-mortems, service restoration
Technologies
Apache Cassandra, Kafka, Kubernetes, relational databases, NoSQL
Responsibilities
Operate and improve the observability platform (metrics, logs, traces, alerting); Monitor critical systems and respond to operational incidents; Lead major incident response and coordinate recovery; Conduct root cause analysis and facilitate post-mortems; Define and monitor SLIs, SLOs, and error budgets; Automate operational processes to reduce toil; Troubleshoot distributed environments; Collaborate with development and infrastructure teams to improve deployment reliability.
Seniority
Senior, hands-on IC