CareerPlanSign in

Staff Site Reliability Engineer - AI Platform Runtime

US, CA, Santa Clara💼 Full-time💰 $168,000–$168,000🗓 2026-09-08 → 2026-09-25

Core

Design, build, and maintain large-scale production systems and AI-driven enterprise products with high efficiency and availability.

Role type

Staff Site Reliability Engineer (AI Platform Runtime)

Builds

Resilient distributed systems, AI Agents, and AI Skills for platform operations

Domain

AI/ML infrastructure, Public Cloud, Systems Architecture

Deliverable

production ML models | infrastructure

Required skills

Site Reliability Engineering, Systems Architecture, Kubernetes, Public Cloud services, Infrastructure-as-Code, Observability, Python, Go, Typescript, JavaScript

Preferred skills

Large-scale automation systems, Technical strategy, Measurable reliability outcomes

Technologies

AWS CDK, AWS CloudFormation, Terraform, CrossPlane, OpenTelemetry, Kubernetes, AWS, Azure, GCP

Responsibilities

Lead technical strategy and roadmap for SRE initiatives, Design and build resilient distributed systems, Architect and develop AI Agents and AI Skills, Drive automation and observability improvements, Collaborate across Cloud, Platform, Security, and AI/ML teams, Analyze and troubleshoot complex systems, Mentor and influence engineers across teams

Seniority

Staff, hands-on IC with strategic influence

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.