Senior Site Reliability Engineer
Core
Ensure system reliability, scalability, and operational excellence for high-availability Cassandra and Elasticsearch clusters in production.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Production-grade distributed data systems (Cassandra, Elasticsearch)
Domain
Cloud-native infrastructure, distributed databases, DevOps
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Cassandra, Elasticsearch, Python, Bash, Terraform, Ansible, CI/CD, Kubernetes, GCP/AWS/Azure
Preferred skills
Prometheus, Grafana, ELK stack, distributed systems design, resiliency patterns
Responsibilities
Maintain and optimize high-availability Cassandra and Elasticsearch clusters; Automate infrastructure tasks, CI/CD pipelines, and operational workflows; Collaborate with development and operations teams to implement SRE best practices; Proactively identify performance bottlenecks and implement preventive measures
Seniority
Senior, hands-on IC