Senior Site Reliability Engineering- CTJ- Secret (Cleared Environments)
Core
Lead reliability and availability improvements, incident response, and automation for multiple Substrate services in highly regulated environments.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Reliability-focused software components, frameworks, and platforms for Substrate services
Domain
Cloud infrastructure, distributed systems, regulated environments (GCCH, DoD, GCCM)
Deliverable
production ML models | infrastructure
Required skills
incident response leadership, post-incident review, automation strategy, monitoring and observability, system architecture influence, mentoring, large-scale cloud or distributed systems experience
Preferred skills
design and implementation of reliability-focused platforms, cross-team coordination
Technologies
cloud platforms, distributed systems, monitoring tools, observability frameworks
Responsibilities
Lead incident response for complex or high-severity incidents, drive high-quality post-incident reviews, drive automation and monitoring strategies, lead design of reliability-focused components, mentor junior SREs, represent SRE perspectives in design reviews
Seniority
Senior, hands-on IC with mentorship responsibilities