Reliability Observability Engineer 2
Core
Lead observability strategy, define SLOs/SLIs, and design scalable monitoring architectures to ensure production readiness and reliability for enterprise applications.
Role type
Senior Reliability Observability Engineer
Builds
Scalable observability architectures, service health dashboards, and actionable alerting models for critical customer journeys.
Domain
Financial services / Distributed systems / Cloud platforms
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Observability Engineering, SRE, SLIs/SLOs/Error Budgets, APM/RUM/Synthetics, Telemetry frameworks, Distributed systems, Microservices, Kubernetes
Preferred skills
Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, OpenTelemetry, Incident analysis/RCA, Stakeholder management
Technologies
Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, OpenTelemetry
Responsibilities
Lead observability strategy across critical customer journeys; Define and govern SLIs, SLOs, and Error Budgets; Design scalable observability architectures; Establish observability governance for dashboards and alerts; Partner with Product, Engineering, and SRE teams; Develop service health dashboards; Analyze telemetry data to reduce alert fatigue.
Seniority
Senior, hands-on IC