Site Reliability Engineer
Core
Ensuring the reliability, performance, and resilience of a cloud-native AI platform processing massive volumes of real-time customer data.
Role type
Site Reliability Engineer (SRE)
Builds
Highly available, fault-tolerant, and scalable cloud infrastructure for an AI-native CX platform.
Domain
Cloud Infrastructure / AI Platform / Customer Experience
Deliverable
infrastructure
Required skills
Kubernetes (EKS, GKE), Infrastructure as Code (Terraform), Cloud platforms (AWS, GCP, Azure), Observability (Prometheus, Grafana, Datadog, ELK), Scripting (Python, Bash), CI/CD pipelines, Distributed systems, Networking, Load balancing
Preferred skills
RabbitMQ, Redis, Ansible, AWX, Multi-cloud environments, Cloud certifications, Linux certifications
Technologies
Terraform, Docker, Kubernetes, Prometheus, Grafana, Datadog, ELK, Jenkins, GitHub Actions, Bitbucket, AWS, GCP, Azure, Python, Bash
Responsibilities
Design and maintain highly available, fault-tolerant, and scalable infrastructure; Proactively identify and eliminate single points of failure; Manage and optimize workloads across cloud environments; Operate and scale Kubernetes clusters; Implement and refine monitoring systems and alerting; Write scripts and build tooling to automate operational work; Collaborate with engineering teams on performance bottlenecks and CI/CD improvements.
Seniority
Mid-level, hands-on IC
