Engineer 2 - Site Reliability Engineering
Core
Operate, improve, and ensure data quality within the monitoring tools landscape for digital media and entertainment services.
Role type
Monitoring Tools Analyst (AIOps group)
Builds
Monitoring toolsets, dashboards, reports, and synthetic tests
Domain
Media and technology / Observability and Reliability
Deliverable
dashboards & analysis
Required skills
Monitoring toolsets (OP5/Nagios, Datadog), Agile DevOps methodologies, CI/CD pipeline & software testing (Ansible, Terraform, Puppet, Chef, Jenkins), large-scale distributed systems troubleshooting, OS management (Windows, Linux), public cloud services (AWS, Azure, Google), network technology concepts (TCP/IP, DNS, SSL, Firewalls)
Preferred skills
Multi-time zone environment experience, analytical mindset, media industry experience, ITSM qualifications
Technologies
OP5, Nagios, Datadog, ServiceNow, Jira, Ansible, Terraform, Puppet, Chef, Jenkins, AWS, Azure, Google
Responsibilities
Responding to and administering events and alerts from all toolsets, testing and managing monitoring software agents, creating and modifying synthetic tests, performing trend analysis of alerts/events, collaborating with incident management and application support teams, creating and maintaining groups, dashboards, and reports
Seniority
Mid-level (2-5 years experience)