Sr Platform Monitoring Engineer
Core
Tech Lead for platform observability, incident response, and proactive monitoring to ensure Databricks platform reliability and customer experience.
Role type
Senior IC Tech Lead (SRE/Platform Engineering)
Builds
End-to-end observability workflows, alerting pipelines, automation tools, and incident response systems for the Databricks Data + AI Platform.
Domain
Cloud Infrastructure & Observability
Deliverable
production ML models | infrastructure
Required skills
Incident lifecycle ownership, root cause analysis, observability architecture, automation tooling, cross-functional coordination, mentorship
Preferred skills
None stated
Technologies
AWS, Azure, GCP, Docker, Kubernetes, ELK, Prometheus, Grafana, PagerDuty, Python
Responsibilities
Lead platform incident investigation and coordination; Conduct post-incident root cause analysis; Design and implement alerting pipelines and observability workflows; Build automation tools and reusable monitoring patterns; Mentor junior engineers on observability patterns; Participate in on-call rotation
Seniority
Senior, hands-on IC with leadership responsibilities