Senior Site Reliability Engineer
Core
Design and maintain scalable, reliable, and secure cloud-native infrastructure while automating operational tasks and ensuring system resilience.
Role type
Senior Site Reliability Engineer
Builds
Production systems and infrastructure for OutSystems' low-code AI development platform
Domain
Cloud infrastructure, SRE, and AI-native software development
Deliverable
production ML models | infrastructure
Required skills
Python, Kubernetes, Linux, Networking, Cloud infrastructure (AWS), Incident response, Automation, SLO/SLA management, Distributed systems debugging
Preferred skills
Go, Bash/Shell scripting, Infrastructure as Code (Terraform, CloudFormation), Monitoring tools (Grafana, Prometheus, ELK), Prompt engineering, AI Native IDEs
Technologies
Python, Go, Kubernetes, EKS, AWS, Terraform, CloudFormation, Grafana, Prometheus, ELK stack
Responsibilities
Lead incident response and conduct root cause analysis; Design and implement scalable, secure infrastructure; Automate operational tasks and incident detection; Establish and maintain SLOs and SLAs; Collaborate with development teams on system resilience; Implement monitoring, alerting, and logging solutions
Seniority
Senior, hands-on IC
