CareerPlanSign in

Reliability Engineer 4 (Observability Specialist )

10 Locations💼 Full-time🗓 2026-08-18 → 2026-09-29

Core

Senior Reliability Engineer establishing observability governance, defining SLIs/SLOs, and ensuring production readiness for enterprise applications.

Role type

Senior IC Reliability Engineer (Observability)

Builds

Production-ready enterprise applications and services with robust observability and reliability metrics.

Domain

Financial Services / Observability Engineering

Required skills

SLI/SLO definition, Observability Governance, Telemetry Analysis, Incident Analysis, Technical Leadership, Distributed Systems Architecture

Preferred skills

APM/RUM/Synthetics expertise, Cloud/Kubernetes proficiency, Alert Governance, RCA

Technologies

Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, OpenTelemetry

Responsibilities

Lead definition and governance of SLIs, SLOs, and error budgets; Design scalable observability architectures and instrumentation standards; Develop executive and operational service health dashboards; Analyze telemetry and incident data to identify gaps and improve detection; Provide technical mentorship on monitoring design patterns and alert governance; Partner with SRE and product teams to ensure applications are fully instrumented.

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 855,000+ jobs from 20+ sources.