Slack Proactive Monitoring Engineer
Core
Proactively monitor and safeguard the health and performance of Slack's largest global enterprise deployments (Enterprise Grid) by detecting anomalies, triaging system exceptions, and orchestrating rapid mitigation steps before customers experience issues.
Role type
Senior IC proactive monitoring engineer (SaaS platform reliability)
Builds
Proactive monitoring operations, automated alerting workflows, and technical health check reviews for enterprise Slack customers
Domain
SaaS platform engineering, Site Reliability Engineering, Customer Success
Deliverable
production ML models | product features | dashboards & analysis | client delivery | infrastructure
Required skills
observability and monitoring tools (Grafana, Splunk, Datadog, PagerDuty), cloud-based SaaS architecture, APIs, log/metrics/trace analysis, root cause analysis, technical advisory, automation scripting
Preferred skills
Slack platform (API, workflows, Bolt framework), Salesforce Service Cloud/OrgCS, Python/JavaScript/Bash scripting, SaaS customer-facing support experience, ITIL/SRE certification
Technologies
Grafana, Splunk, Datadog, PagerDuty, Slack API, Slack Workflow Builder, Enterprise Key Management, Identity Provider
Responsibilities
Continuously monitor dashboards, alerting systems, and telemetry data for early signals of degradation; Triage and correlate alerts from multiple sources to identify patterns before customers report issues; Actively monitor Slack platform health dashboards, network latency signals, message delivery queues, and database capacities; Monitor critical custom automations, Slack Workflow Builder runs, Enterprise Key Management (EKM) operations, and Identity Provider (IDP) authentication syncs; Identify customers potentially affected by degraded service conditions and coordinate proactive outreach with Customer Success and Support teams; Partner with the Incident Management team to escalate signals that meet incident-threshold criteria; Perform root cause analysis (RCA) on proactively detected issues; Work closely with Engineering and SRE teams to drive rapid remediation of identified issues; Intervene in low-risk system exceptions (e.g., advising clients on misconfigured Slack Webhooks, API rate limit exhaustion, or broken Salesforce-Slack app connections); Build and maintain Slack-based automations and workflows to streamline proactive monitoring operations
Seniority
Senior, hands-on IC