Site Reliability Engineer - Devops, AWS, Python/GO - 4 to 8 years
Core
Design, build, and optimize cloud and data infrastructure to ensure high availability, reliability, and scalability of SaaS systems operating at multi-region scale.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Cloud and data platform infrastructure, automation tooling, and monitoring systems
Domain
Cloud Infrastructure / SaaS / Data Platforms
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes (EKS/self-managed), Docker, AWS, Infrastructure as Code (Terraform/Ansible/CloudFormation), Observability (Prometheus/Grafana/CloudWatch/Splunk/Open Telemetry), Python/Go/Bash, Unix/Linux systems
Preferred skills
Cloud-based data platform management, AI-assisted SRE tooling, automation-first approaches
Technologies
Kubernetes, Docker, AWS, Terraform, Ansible, CloudFormation, Prometheus, Grafana, CloudWatch, Splunk, Open Telemetry, Python, Go, Bash
Responsibilities
Design and optimize cloud infrastructure for high availability; Collaborate with cross-functional teams to create scalable solutions; Monitor production systems and manage on-call rotations; Conduct root cause analysis and postmortem reviews; Build automation to integrate software and systems engineering
Seniority
Senior, hands-on IC