CareerPlanSign in

Staff Service Reliability and Operational Intelligence Engineer

Santa Clara, California🌐 Remote💼 Full-time🗓 2026-09-17 → 2026-09-26

Core

Define technical direction for production reliability, observability architecture, and AIOps workflows to ensure service continuity and resilience for cloud-managed SaaS products.

Role type

Staff Service Reliability and Operational Intelligence Engineer

Builds

Scalable infrastructure for cloud-managed SaaS products with on-premises components

Domain

Cloud Infrastructure / SaaS / Observability / AIOps

Deliverable

production ML models | infrastructure

Required skills

Distributed systems, Kubernetes, AWS/GCP, CI/CD, Python/Go, Infrastructure as Code, Observability architecture, Incident command, Capacity planning, Technical leadership

Preferred skills

AI Ops design, Autonomous remediation, Amazon Bedrock AgentCore, FinOps, Load balancing, Health-based failover

Technologies

AWS, GCP, Kubernetes, Python, Go, Jira, Confluence, Amazon Bedrock AgentCore

Responsibilities

Shape multi-year roadmap for operational excellence; Define New Service Introduction framework; Establish service ownership standards; Lead observability platform architecture; Govern reliability model (SLIs/SLOs); Advance incident management maturity; Design AIOps capabilities; Improve on-call effectiveness; Provide hands-on leadership during major incidents.

Seniority

Staff, hands-on IC with strategic scope

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.