Principal Site Reliability Engineering Manager - CTJ -Secret
Core
Lead and develop a team of Site Reliability Engineers to own the operational health and reliability of Substrate services in regulated environments.
Role type
Principal Site Reliability Engineering Manager (hands-on IC)
Builds
Substrate services running in regulated environments
Domain
Cloud infrastructure, regulated environments, disaster recovery
Deliverable
production ML models | infrastructure
Required skills
Team leadership and coaching, incident management and post-incident reviews, SLO/SLI definition and monitoring, disaster recovery planning and execution, automation and toil reduction, cross-functional partnership, strategic influence, compliance and security integration
Preferred skills
AI-assisted operational techniques, regulated environment operations
Technologies
Large-scale cloud or distributed systems
Responsibilities
Provide clear expectations, regular coaching, and career guidance for senior and principal SREs; Drive change and influence across the org to establish and drive SLOs, SLIs, and operational metrics; Lead effective incident management and post-incident reviews emphasizing systemic fixes; Serve as an actively engaged on-call engineer leading incident response; Own reliability, resilience, and disaster recovery including DR and game day exercises; Partner with engineering and product teams to embed reliability, security, and compliance considerations early; Represent team work to leadership and partners articulating risks and progress; Build strong cross-functional relationships with security, compliance, and operations teams