Staff Site Reliability Engineer
Core
Set technical direction for reliability across regions and services, own reliability strategy, and lead high-severity incident response for quantum computing platforms.
Role type
Staff Site Reliability Engineer
Builds
Quantum computing platforms (computing, networking, sensing, security) available via major cloud providers
Domain
Quantum Computing / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Large-scale fault-tolerant system operations, observability stack design, SLO and error budget governance, chaos engineering, incident command, disaster recovery testing, cloud security posture management, capacity planning, AI Ops workflows, multi-team technical leadership
Preferred skills
Cloud security posture management, risk prioritization using identity and exposure context, autonomous remediation and self-healing workflows, capacity management and FinOps, load-balancing design, AI traffic management via LLM gateway
Technologies
AWS, GCP, Amazon Bedrock Agent Core, LLM gateway
Responsibilities
Own service-level objectives and production reliability outcomes, design and operate observability stacks, define and manage SLOs and error budgets, design and execute chaos experiments, serve as incident commander for highest-severity incidents, manage on-call rotations and escalation paths, own disaster-recovery testing and failover validation, co-own cloud security posture management, own reliability of stateful and streaming services, mentor engineers and set cross-team standards
Seniority
Staff, hands-on IC with strategic leadership