Staff Service Reliability and Operational Intelligence Engineer
Core
Define technical direction for production reliability, observability architecture, and AIOps workflows to ensure service continuity and resilience for cloud-managed SaaS products.
Role type
Staff Service Reliability and Operational Intelligence Engineer
Builds
Scalable infrastructure for cloud-managed SaaS products with on-premises components
Domain
Cloud Infrastructure / SaaS / Observability / AIOps
Deliverable
production ML models | infrastructure
Required skills
Distributed systems, Kubernetes, AWS/GCP, CI/CD, Python/Go, Infrastructure as Code, Observability architecture, Incident command, Capacity planning, Technical leadership
Preferred skills
AI Ops design, Autonomous remediation, Amazon Bedrock AgentCore, FinOps, Load balancing, Health-based failover
Technologies
AWS, GCP, Kubernetes, Python, Go, Jira, Confluence, Amazon Bedrock AgentCore
Responsibilities
Shape multi-year roadmap for operational excellence; Define New Service Introduction framework; Establish service ownership standards; Lead observability platform architecture; Govern reliability model (SLIs/SLOs); Advance incident management maturity; Design AIOps capabilities; Improve on-call effectiveness; Provide hands-on leadership during major incidents.
Seniority
Staff, hands-on IC with strategic scope