CareerPlanGet AI match score →

Senior Site Reliability Engineer

US - San Francisco Bay Area💼 Full-time🗓 2026-06-02 → 2026-07-31

Core

Design and maintain scalable, reliable, and secure cloud-native infrastructure while automating operational tasks and ensuring system resilience.

Role type

Senior Site Reliability Engineer

Builds

Production systems and infrastructure for OutSystems' low-code AI development platform

Domain

Cloud infrastructure, SRE, and AI-native software development

Deliverable

production ML models | infrastructure

Required skills

Python, Kubernetes, Linux, Networking, Cloud infrastructure (AWS), Incident response, Automation, SLO/SLA management, Distributed systems debugging

Preferred skills

Go, Bash/Shell scripting, Infrastructure as Code (Terraform, CloudFormation), Monitoring tools (Grafana, Prometheus, ELK), Prompt engineering, AI Native IDEs

Technologies

Python, Go, Kubernetes, EKS, AWS, Terraform, CloudFormation, Grafana, Prometheus, ELK stack

Responsibilities

Lead incident response and conduct root cause analysis; Design and implement scalable, secure infrastructure; Automate operational tasks and incident detection; Establish and maintain SLOs and SLAs; Collaborate with development teams on system resilience; Implement monitoring, alerting, and logging solutions

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗