Expert Observability Engineer
Core
Architect and govern a unified enterprise observability framework for metrics, logs, traces, and events, leading FOAK implementations and establishing reliability practices.
Role type
Senior IC Principal Observability Engineer
Builds
Secure, repeatable production patterns for telemetry, data retention, and observability cost management across Kubernetes, Docker, and multi-cloud environments.
Domain
Cloud Infrastructure & Observability
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Docker, OpenShift, AWS, Azure, GCP, Linux/RHEL, Windows Server, VMware, Citrix VDI, load balancers, edge proxies, Ansible, Terraform, Python, Bash, CI/CD tools, ServiceNow, ITIL 4, major incident management, scalable telemetry pipelines, time-series databases, FOAK rollouts, vendor transition programs.
Preferred skills
CKA, AWS or Azure Cloud Architect, APM or observability vendor certifications.
Technologies
IBM Instana, Grafana Enterprise, Alloy, Prometheus, OpenTelemetry, Telegraf, InfluxDB, SolarWinds, Netcool, Elastic, Splunk, GitHub Actions, GitLab, Jenkins, SulAmérica, Prudential, MetLife.
Responsibilities
Architect and govern enterprise observability framework; Lead FOAK observability implementations; Define telemetry, data-retention, and cost-management standards; Establish SLIs, SLOs, error budgets, and reliability practices; Lead P1/P2 incident war rooms and root cause analyses; Reduce detection and recovery times through event correlation and dynamic thresholds; Drive Observability-as-Code and infrastructure automation; Design observability for multi-cloud environments; Integrate observability platforms with ITSM and delivery pipelines; Lead knowledge-transfer programs and technical mentoring.
Seniority
Senior, hands-on IC