Expert Automation & Observability Engineer
Core
Architect and govern a unified enterprise observability framework (metrics, logs, traces, events) to transform decentralized monitoring into a proactive, automated system for hybrid cloud and legacy environments.
Role type
Senior IC SRE/Principal Observability Architect
Builds
Scalable observability platforms, automated telemetry pipelines, and SRE practices for enterprise clients
Domain
Cloud-native infrastructure, Site Reliability Engineering, Observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Enterprise observability architecture, SRE practices (SLIs/SLOs), FOAK technology implementation, Infrastructure automation (Ansible/Terraform), Cloud-native Kubernetes observability, Incident management & RCA, Knowledge transfer leadership
Preferred skills
CKA certification, Cloud Architect certification, APM/Observability vendor certifications
Technologies
IBM Instana, Grafana, OpenTelemetry, Telegraf, InfluxDB, Prometheus, SolarWinds, Netcool, Elastic/Splunk, Kubernetes, Docker, OpenShift, AWS, Azure, GCP, Linux, Windows Server, VMware, ServiceNow, ITIL 4, GitHub Actions, GitLab, Jenkins
Responsibilities
Architect unified observability frameworks; Define SLIs/SLOs and error budgets; Lead P1/P2 incident war rooms and RCA; Automate deployment and self-healing workflows; Design deep observability for Kubernetes and multi-cloud; Lead knowledge transfer and vendor transition programs
Seniority
Senior, hands-on IC with strategic leadership