CareerPlanGet AI match score →

Incident Management Reliability Engineer

Budapest💼 Full-time💰 $15,680,000–$15,680,000🗓 2026-06-24 → 2026-07-30

Core

Ensuring the stability, resilience, and reliability of critical IT services by leading major incident management and implementing reliability engineering principles.

Role type

Senior Incident Management Reliability Engineer

Builds

Stable, resilient IT services with minimized disruptions and rapid recovery capabilities.

Domain

IT Operations / Service Reliability

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Incident management, Root cause analysis, Post-incident review facilitation, SLO/SLI/SLA definition, Capacity planning, Change review participation, Automation scripting, Observability implementation, Stakeholder communication, ITSM best practices

Preferred skills

ITIL v4, SRE Foundation, Cloud certifications, Incident Command System

Technologies

AIOps tools, Runbooks, Cloud platforms, Containerization, Infrastructure as code

Responsibilities

Lead end-to-end management of Major Incidents (P1/P2) and act as command centre lead during critical outages; Collaborate with service owners to enhance service reliability, observability, and fault tolerance; Implement proactive monitoring, alerting, and automated recovery mechanisms; Develop and track SLOs/SLIs/SLAs to measure service health; Automate operational tasks and enhance service recovery using scripts and AIOps tools.

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗