CareerPlanSign in

Incident Manager

🌐 Remote💼 Full-time💰 $103,900–$103,900🗓 2026-06-30 → 2026-09-25

Core

Lead critical production incidents for Databricks' cloud-based data and AI infrastructure, orchestrating multi-team responses and ensuring customer/stakeholder confidence during high-impact events.

Role type

Senior Incident Manager (SRE)

Builds

Cloud-native data and AI infrastructure services (Databricks Data Intelligence Platform)

Domain

Cloud Infrastructure / Data & AI

Deliverable

production ML models | infrastructure

Required skills

Incident command & coordination, technical root cause analysis, cloud infrastructure knowledge (AWS/Azure/GCP), log analysis & debugging, observability systems (metrics/logging/tracing), scripting (Python/Go/Bash), incident playbook development, executive & customer communication

Preferred skills

Experience with distributed systems, mentoring peers in incident response

Technologies

Datadog, Elasticsearch, Splunk, Cloud Logging, OpenTelemetry, Prometheus, Grafana

Responsibilities

Coordinate multi-disciplinary response efforts to mitigate impact and restore operations; trace and document underlying causes across distributed systems; deliver frequent updates to internal stakeholders and publish customer-facing notifications; summarize key learnings and ensure procedural improvements are followed through; mentor peers in incident communication and technical response disciplines

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.