Principal Product Manager, Reliability Platform (Observability, SRE, AIM)
Core
Drive the core reliability platforms and services that empower engineering teams to deliver software quickly, safely, and with high quality.
Role type
Principal Product Manager, Reliability Platform (Observability, SRE, AIM)
Builds
Internal developer engineering tools and frameworks for system availability, incident management, and observability.
Domain
Insurance technology / Developer Platform Engineering / Site Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Technical product management, system availability strategy, incident management, cloud infrastructure, observability, data-driven decision making, agile environment leadership, developer workflow optimization, operational health metrics, KPI definition
Preferred skills
MBA, internal developer platforms (IDPs), service catalogs, Paved Road engineering, modern observability tools (Grafana), cloud platforms (Azure, AWS), container orchestration (Kubernetes), SRE principles (SLOs, error budgets)
Technologies
Grafana, Azure, AWS, Kubernetes
Responsibilities
Build and scale foundational developer platforms; Define and execute product strategy for Observability, BCDR & Incident Management; Lead cross-functional teams to deliver developer-facing products; Deeply understand the developer workflow to remove friction; Own and prioritize the product roadmap for platform services; Define and track operational health metrics like system availability and MTTR; Champion a culture of reliability and ownership; Identify and measure KPIs reflecting developer productivity and system health
Seniority
Principal, hands-on IC with strategy & mentorship