Incident Manager
Core
Lead critical production incidents for Databricks' cloud-based data and AI infrastructure, orchestrating multi-team responses and ensuring customer/stakeholder confidence during high-impact events.
Role type
Senior Incident Manager (SRE)
Builds
Cloud-native data and AI infrastructure services (Databricks Data Intelligence Platform)
Domain
Cloud Infrastructure / Data & AI
Deliverable
production ML models | infrastructure
Required skills
Incident command & coordination, technical root cause analysis, cloud infrastructure knowledge (AWS/Azure/GCP), log analysis & debugging, observability systems (metrics/logging/tracing), scripting (Python/Go/Bash), incident playbook development, executive & customer communication
Preferred skills
Experience with distributed systems, mentoring peers in incident response
Technologies
Datadog, Elasticsearch, Splunk, Cloud Logging, OpenTelemetry, Prometheus, Grafana
Responsibilities
Coordinate multi-disciplinary response efforts to mitigate impact and restore operations; trace and document underlying causes across distributed systems; deliver frequent updates to internal stakeholders and publish customer-facing notifications; summarize key learnings and ensure procedural improvements are followed through; mentor peers in incident communication and technical response disciplines
Seniority
Senior, hands-on IC