Senior VP, Site Reliability Engineer
Core
Design and implement end-to-end observability, automate operational work, and ensure reliability of distributed financial systems.
Role type
Senior VP, Site Reliability Engineer (hands-on IC with strategic scope)
Builds
Production observability frameworks, self-healing solutions, and automated remediation tooling for distributed systems.
Domain
Financial services / Distributed systems / Observability
Required skills
Java, Python, Bash, distributed systems, microservices, CI/CD, cloud platforms, Kubernetes, observability platforms (AppDynamics, Dynatrace, Grafana, Splunk), incident management, SLO/SLI definition, capacity planning, performance tuning
Preferred skills
SRE concepts, automation scripting, system architecture input
Technologies
Bash, CI/CD, Cloud, Dynatrace, Grafana, Java, Kubernetes, Python, Splunk
Responsibilities
Design and implement end-to-end observability across distributed systems; integrate and optimize monitoring tools; develop dashboards, alerts, and telemetry frameworks; automate repetitive operational work; build self-healing and auto-remediation solutions; troubleshoot and resolve complex production issues; participate in incident management and root cause analysis; define and measure service health using SLIs, SLOs, and key performance metrics; contribute to performance tuning and capacity planning; provide input on system architecture to improve resilience and scalability.
Seniority
Senior VP, hands-on IC with strategic oversight


