Senior Site Reliability Engineer – Unified Observability
Core
Lead the design, implementation, and operational maturity of a unified enterprise observability platform delivering end-to-end visibility across Restaurants, Retail, and Payments environments.
Role type
Senior Site Reliability Engineer (Unified Observability)
Builds
Unified enterprise observability platform for NCR Voyix Restaurants, Retail, and Payments
Domain
Unified Commerce / Cloud Operations / Observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes (AKS, GKE), Azure, Google Cloud Platform, Grafana, Datadog, Prometheus, OpenTelemetry, Dynatrace, New Relic, Terraform, Python, Go, PowerShell, SLIs, SLOs, error budgets, CI/CD integration, automation, incident response, mentorship
Preferred skills
AI Ops, event correlation, predictive monitoring, ServiceNow integrations, ITSM/ITOM processes, cloud certifications, technical lead experience
Technologies
Azure, Google Cloud Platform, Kubernetes, AKS, GKE, Grafana, Datadog, Prometheus, OpenTelemetry, Dynatrace, New Relic, ServiceNow, Terraform
Responsibilities
Lead architecture and design of enterprise observability solutions; Establish enterprise observability standards for monitoring, logging, and tracing; Develop executive and engineering dashboards; Define reliability frameworks including SLIs, SLOs, and error budgets; Partner cross-functionally to identify reliability risks; Improve incident prevention and recovery through automation; Integrate observability with CI/CD and operational workflows; Influence technical strategy and roadmap; Mentor engineers as subject matter expert; Establish governance and best practices across teams
Seniority
Senior, hands-on IC with leadership and mentorship responsibilities