Senior Site Reliability Engineer
Core
Define and drive reliability, scalability, and security architecture for a mission-critical low-code/AI development platform.
Role type
Senior Site Reliability Engineer (Technical Lead)
Builds
Resilient cloud-native infrastructure, automation frameworks, and observability solutions for enterprise AI applications.
Domain
Enterprise Software / Low-Code Development / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Linux internals, cloud networking, Infrastructure as Code (Terraform), Python, Go, incident response, SLO/SLI design, automation tooling
Preferred skills
Kubernetes certifications (CKA, CKS), cloud provider certifications (AWS Solutions Architect), compliance frameworks (SOC 2, CIS), Gen AI tooling
Technologies
Kubernetes, Terraform, AWS, Python, Go
Responsibilities
Architect SLO/SLI frameworks and error-budget policies; Lead design and rollout of IaC and Kubernetes environments; Drive observability strategy for distributed systems; Act as Incident Commander during critical events; Build core automation frameworks and platform tools; Mentor software and reliability engineers
Seniority
Senior, hands-on IC with mentorship responsibilities