CareerPlanGet AI match score →

Platform Monitoring & Incident Engineer

Chicago💼 Full-time💰 $100,000–$100,000🗓 2026-07-13 → 2026-07-31

Core

On-call monitoring of a global payments platform, coordinating high-impact incidents, and driving proactive reliability improvements for merchants.

Role type

Senior Platform Monitoring & Incident Engineer

Builds

A global payments platform serving merchants like Meta, Uber, and H&M

Domain

Financial Technology / Payments

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Incident management, problem management, root cause analysis, on-call rotation, alert investigation, automation, stakeholder communication, data visualization, process standardization

Preferred skills

Experience with monitoring frameworks, merchant impact prioritization, cross-functional collaboration

Technologies

Prometheus, Grafana, ELK Stack, Datadog, Dynatrace, Splunk

Responsibilities

Observe platform and merchant performance to detect issues proactively; coordinate mitigation and recovery of high-impact incidents; communicate real-time status to merchants during incidents; analyze incident trends to identify recurring issues and drive long-term fixes; develop automation for effective monitoring; document learnings in the monitoring playbook.

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗