Staff Site Reliability Engineer - AI Platform Runtime
Core
Design, build, and maintain large-scale production systems and AI-driven enterprise products with high efficiency and availability.
Role type
Staff Site Reliability Engineer (AI Platform Runtime)
Builds
Resilient distributed systems, AI Agents, and AI Skills for platform operations
Domain
AI/ML infrastructure, Public Cloud, Systems Architecture
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering, Systems Architecture, Kubernetes, Public Cloud services, Infrastructure-as-Code, Observability, Python, Go, Typescript, JavaScript
Preferred skills
Large-scale automation systems, Technical strategy, Measurable reliability outcomes
Technologies
AWS CDK, AWS CloudFormation, Terraform, CrossPlane, OpenTelemetry, Kubernetes, AWS, Azure, GCP
Responsibilities
Lead technical strategy and roadmap for SRE initiatives, Design and build resilient distributed systems, Architect and develop AI Agents and AI Skills, Drive automation and observability improvements, Collaborate across Cloud, Platform, Security, and AI/ML teams, Analyze and troubleshoot complex systems, Mentor and influence engineers across teams
Seniority
Staff, hands-on IC with strategic influence
