Senior Automation & Observability Engineer
Core
Design, implement, and maintain enterprise monitoring and observability solutions for infrastructure, applications, IoT platforms, and telemetry ecosystems to ensure high availability and reliability.
Role type
Senior IC Automation & Observability Engineer
Builds
Production monitoring dashboards, alerting systems, automation frameworks, and telemetry pipelines
Domain
Enterprise IT Operations, Cloud Infrastructure, IoT, Observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Grafana, IBM Instana, SolarWinds, Telegraf, Prometheus, InfluxDB, OpenTelemetry, Python, PowerShell, Bash, Ansible, Puppet, Linux Administration, Windows Server, VMware, Incident Management, RCA, SLO/SLA Monitoring
Preferred skills
Docker, Kubernetes, AWS, Azure, GCP, Jenkins, CI/CD pipelines, Microservices Monitoring
Technologies
Grafana, IBM Instana, SolarWinds, Telegraf, Prometheus, InfluxDB, OpenTelemetry, Grafana Alloy, Ansible, Puppet, ServiceNow, Python, PowerShell, Bash, VBScript
Responsibilities
Design and maintain enterprise monitoring solutions; Develop dashboards and visualizations; Monitor infrastructure, applications, and IoT services; Perform Root Cause Analysis and troubleshooting; Automate operational tasks and remediation workflows; Manage incident response and major incident bridges; Configure and maintain time-series databases and telemetry pipelines.
Seniority
Senior, hands-on IC