Site Reliability Engineer II - CTJ - Secret
Core
Own reliability and operational health for Substrate components/services in highly regulated environments, acting as an on-call engineer to diagnose and resolve production incidents.
Role type
Site Reliability Engineer II
Builds
Automation to reduce operational toil and improve service stability
Domain
Cloud infrastructure, regulated government environments
Deliverable
production ML models | infrastructure
Required skills
incident response, root cause analysis, automation design, monitoring and alerting, SLO management, post-incident reviews, collaboration with software engineering teams
Preferred skills
experience with large-scale cloud or distributed systems, experience in regulated/sovereign/compliance-sensitive environments
Technologies
cloud platforms, distributed systems, monitoring tools, automation frameworks
Responsibilities
Participate in on-call rotation and independently respond to incidents; Design and implement automation to improve reliability and efficiency; Develop and maintain monitoring, alerting, and telemetry to support SLOs; Lead post-incident reviews focusing on root cause analysis; Collaborate with engineering teams to embed reliability into service design