Principal Site Reliability Engineer
Core
Define enterprise standards for end-to-end service ownership, lead cross-functional initiatives shaping system architecture, and set company strategy for observability and governance of SLIs/SLOs.
Role type
Principal Site Reliability Engineer (Architect level)
Builds
Enterprise digital systems, specialized AI, and data-driven foundation for insurance agencies, MGAs, and carriers
Domain
Insurance technology, Cloud Infrastructure, Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
C#, .NET, Java, Python, React, AWS, Kubernetes, CI/CD, Infrastructure as Code, Linux, Windows, relational databases, SLIs, SLOs, error budgets
Preferred skills
Insurance industry exposure
Technologies
AWS, Kubernetes, CI/CD, Linux, Windows, ASP.NET, ARM
Responsibilities
Define enterprise standards for end-to-end service ownership, lead cross-functional initiatives shaping system architecture, set company strategy for observability, establish governance for SLIs and SLOs, advocate for error budgets, lead incident response for critical high-severity events, promote blameless postmortem culture
Seniority
Principal, hands-on IC with strategic leadership