Sr Engineer - Applications & Observability
Core
Design, implement, and manage monitoring and observability solutions for servers, applications, databases, network devices, and cloud environments to ensure SLA adherence and system health visibility.
Role type
Senior IC applications & observability engineer
Builds
Monitoring platforms, dashboards, alerting mechanisms, and operational insights for enterprise clients
Domain
Cloud infrastructure, enterprise applications, and IT operations
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Observability platform configuration, metrics/logs/traces analysis, alerting design, root cause analysis, capacity planning, scripting/automation, Infrastructure as Code, CI/CD integration, cloud monitoring services
Preferred skills
AIOps, self-healing operations, event correlation, dependency mapping, cloud migration support
Technologies
Datadog, Dynatrace, New Relic, Splunk, Elastic, Prometheus, Grafana, Terraform, Ansible, CloudFormation, AWS, Azure, GCP, Python, PowerShell, Bash
Responsibilities
Design and implement observability solutions for multi-environment infrastructure; Configure and maintain observability platforms for metrics, logs, traces, and alerts; Develop dashboards and reports for system health and operational KPIs; Set up and fine-tune alerting mechanisms and notification workflows; Perform analysis of telemetry data to identify trends, anomalies, and service risks; Integrate monitoring tools with ITSM, event management, and ticketing platforms; Automate monitoring configuration and deployment using scripting and Infrastructure as Code practices.
Seniority
Senior, hands-on IC