Senior Site Reliability Engineer
Core
Building intelligent monitoring, alerting, reliability testing, and automated remediation for critical wealth management applications and platforms.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Scalable SRE solutions, modern observability practices, and AI-enhanced self-healing operations for internal users and clients.
Domain
Financial services / Wealth Management / Cloud-native operations
Deliverable
production ML models | infrastructure
Required skills
Infrastructure automation (Ansible), Scripting (Bash, Python, PowerShell), Observability tooling (Elasticsearch, Dynatrace, Kubernetes, OpenShift, Kafka), Incident management, SLI/SLO definition, Cloud-native distributed systems concepts, AIOps/AI/ML concepts.
Preferred skills
Financial services domain experience, OpenTelemetry, Prometheus/Grafana/Splunk, CI/CD tools (Jenkins, Artifactory, Vault), Containerization (Docker), Anomaly detection solutions, AI governance.
Technologies
Elasticsearch, Ansible, GitHub Actions, Dynatrace, PagerDuty, Moogsoft, Kubernetes, OpenShift, Kafka, Bash, Python, PowerShell, OpenTelemetry, Prometheus, Grafana, Splunk, Jenkins, Artifactory, Vault, Docker
Responsibilities
Develop intelligent monitoring, alerting, and automated remediation capabilities; Deploy metrics, logs, traces, and actionable alerting; Design ML-based anomaly detection and self-healing solutions; Standardize telemetry and instrumentation; Automate operational workflows using scripting and configuration management; Define and track service health metrics (SLIs, SLOs, error budgets); Partner with development teams to ensure reliability standards; Lead incident and problem management including root cause analysis; Drive continuous improvement of operations using engineering and AI-driven approaches.
Seniority
Senior, hands-on IC