Expert Observability Engineer
Core
Design and govern unified enterprise observability strategies spanning metrics, logs, traces, and events to improve service reliability across complex technology landscapes.
Role type
Senior IC Expert Observability Engineer (SRE/Platform)
Builds
Secure, scalable, production-ready observability patterns and telemetry pipelines for global engineering and operations teams.
Domain
Cloud-native monitoring, SRE, and enterprise observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Unified observability architecture, SRE practices, incident management, infrastructure-as-code, telemetry pipeline design, SLO/SLI management, FOAK implementation, vendor transition leadership, automation mindset, root cause analysis, distributed systems expertise
Preferred skills
CKA certification, cloud architecture certifications (AWS/Azure), APM/observability certifications
Technologies
IBM Instana, Grafana, OpenTelemetry, Telegraf, InfluxDB, Prometheus, SolarWinds, Netcool, Elastic, Splunk, Kubernetes, Docker, OpenShift, AWS, Azure, GCP, Linux/RHEL, Windows Server, VMware, Ansible, Terraform, Python, Bash, ServiceNow, GitHub Actions, GitLab, Jenkins
Responsibilities
Architect and govern a unified enterprise observability framework covering metrics, logs, traces, and events; Lead First-of-a-Kind (FOAK) implementations and convert emerging technologies into secure production patterns; Define enterprise standards for telemetry pipelines, data retention, and observability cost optimization; Establish and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets; Act as a senior technical escalation point for critical operational issues and major incidents; Lead P1/P2 incident war rooms and drive evidence-based Root Cause Analysis (RCA); Reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through event correlation and dynamic thresholds; Drive Observability-as-Code and infrastructure automation; Integrate observability platforms with ITSM tools and CI/CD pipelines; Design deep observability capabilities across Docker, Kubernetes, and multi-cloud environments; Correlate application performance telemetry with Kubernetes control planes and infrastructure dependencies; Establish secure-by-design telemetry pipelines; Lead complex knowledge-transfer programs and mentor cross-functional teams; Contribute to enterprise technology roadmaps and influence architectural decisions; Modernize legacy monitoring environments
Seniority
Senior, hands-on IC with leadership responsibilities