CareerPlanSign in

Principal Site Reliability Engineering Manager - CTJ -Secret

United States, Multiple Locations, Multiple Locations💼 Full-time🗓 2026-08-27 → 2026-09-26

Core

Lead and develop a team of Site Reliability Engineers to own the operational health and reliability of Substrate services in regulated environments.

Role type

Principal Site Reliability Engineering Manager (hands-on IC)

Builds

Substrate services running in regulated environments

Domain

Cloud infrastructure, regulated environments, disaster recovery

Deliverable

production ML models | infrastructure

Required skills

Team leadership and coaching, incident management and post-incident reviews, SLO/SLI definition and monitoring, disaster recovery planning and execution, automation and toil reduction, cross-functional partnership, strategic influence, compliance and security integration

Preferred skills

AI-assisted operational techniques, regulated environment operations

Technologies

Large-scale cloud or distributed systems

Responsibilities

Provide clear expectations, regular coaching, and career guidance for senior and principal SREs; Drive change and influence across the org to establish and drive SLOs, SLIs, and operational metrics; Lead effective incident management and post-incident reviews emphasizing systemic fixes; Serve as an actively engaged on-call engineer leading incident response; Own reliability, resilience, and disaster recovery including DR and game day exercises; Partner with engineering and product teams to embed reliability, security, and compliance considerations early; Represent team work to leadership and partners articulating risks and progress; Build strong cross-functional relationships with security, compliance, and operations teams

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.