Site Reliability Engineer
Core
Ensure reliability, scalability, and security of cloud platforms and SaaS systems operating at multi-region scale.
Role type
Site Reliability Engineer (SRE)
Builds
Cloud and data infrastructure, automation tools, and core data platform capabilities
Domain
Cloud infrastructure and SaaS platforms
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes (EKS or self-managed), Docker, cloud platforms (AWS), programming/scripting (Python, Bash, Go), Infrastructure as Code (Ansible, Terraform), Linux systems, CI/CD pipelines, observability tools (ELK, Splunk, Prometheus)
Preferred skills
Cloud-based data platform management, AI-assisted SRE tooling, automation-first approaches
Technologies
Kubernetes, Docker, AWS, Python, Bash, Go, Ansible, Terraform, ELK, Splunk, Prometheus
Responsibilities
Build, deploy, and optimize cloud infrastructure; integrate software and systems engineering; monitor production systems and troubleshoot incidents; contribute to root cause analysis and postmortem reviews; collaborate with cross-functional teams to enhance operational efficiency through automation
Seniority
Mid-level, hands-on IC