CareerPlanSign in

Site Reliability Engineer II - CTJ - Secret

United States, Multiple Locations, Multiple Locations💼 Full-time🗓 2026-08-28 → 2026-09-26

Core

Own reliability and operational health for Substrate components/services in highly regulated environments, acting as an on-call engineer to diagnose and resolve production incidents.

Role type

Site Reliability Engineer II

Builds

Automation to reduce operational toil and improve service stability

Domain

Cloud infrastructure, regulated government environments

Deliverable

production ML models | infrastructure

Required skills

incident response, root cause analysis, automation design, monitoring and alerting, SLO management, post-incident reviews, collaboration with software engineering teams

Preferred skills

experience with large-scale cloud or distributed systems, experience in regulated/sovereign/compliance-sensitive environments

Technologies

cloud platforms, distributed systems, monitoring tools, automation frameworks

Responsibilities

Participate in on-call rotation and independently respond to incidents; Design and implement automation to improve reliability and efficiency; Develop and maintain monitoring, alerting, and telemetry to support SLOs; Lead post-incident reviews focusing on root cause analysis; Collaborate with engineering teams to embed reliability into service design

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.