Staff Site Reliability Engineer
Core
Setting technical direction for reliability across regions and services for cloud-managed SaaS products with on-premises components, focusing on service continuity and customer experience.
Role type
Staff Site Reliability Engineer (Platform Engineering)
Builds
Scalable infrastructure for cloud-managed SaaS products with on-premises components deployed at customer sites
Domain
Quantum computing platform infrastructure, Cloud Operations, SaaS
Deliverable
production ML models | infrastructure
Required skills
Production reliability strategy, Observability architecture, SLO and error budget governance, Chaos engineering, Incident command, Disaster recovery planning, Cloud security posture management, Stateful and streaming platform reliability, AI Ops workflows, Multi-team technical leadership
Preferred skills
Cloud security posture management (CSPM), Runtime vulnerability detection, Autonomous remediation and self-healing (AIOps), Capacity management and FinOps, Load-balancing design, AI traffic management via LLM gateway
Technologies
AWS, GCP, Postgres, Redis/Valkey, Kafka, OpenSearch, Amazon Bedrock Agent Core
Responsibilities
Own service-level objectives, error budgets, and production reliability outcomes end to end; Design and operate the observability stack; Define and manage service-level objectives and drive corrective action; Design and execute chaos experiments; Serve as incident commander for highest-severity incidents; Establish and manage on-call rotations and escalation paths; Own disaster-recovery testing and failover validation; Co-own cloud security posture management, runtime vulnerability detection, and configuration-compliance monitoring; Own reliability of stateful and streaming services, capacity planning, and autonomous agents for triage and self-healing; Mentor engineers and align teams behind a shared reliability roadmap
Seniority
Staff, hands-on IC with strategic leadership