CareerPlanGet AI match score →

Senior Site Reliability Engineer

India💼 Full-time🗓 2026-06-08 → 2026-07-31

Core

Own the reliability of distributed data systems (streaming runtime and processing engines) moving hundreds of billions of rows per day for top-tier enterprises.

Role type

Senior Site Reliability Engineer (Big Data Stack)

Builds

Streaming runtime, processing engines, stateful services, and Kubernetes clusters for data movement.

Domain

Data Integration / Big Data Infrastructure

Deliverable

infrastructure

Required skills

Kafka operations, Kubernetes (EKS/GKE), Redis clustering, Terraform, Python, distributed systems debugging, incident management, capacity planning, observability.

Preferred skills

Spark, Flink, Ray, JVM debugging, CI/CD, AWS services, Ansible.

Technologies

Kafka, Strimzi, Spark, Flink, Ray, Redis, Kubernetes, EKS, GKE, Terraform, Python, Snowflake, BigQuery, Jenkins, GitHub Actions, GitLab CI.

Responsibilities

Manage Kafka broker health, topic lifecycle, and partition tuning at massive scale; operate and tune distributed processing workloads (batch/streaming); run Redis clusters with failover and persistence; manage Kubernetes clusters and operators; build data-aware monitoring and observability; lead root-cause analysis for distributed system failures; provision infrastructure via Terraform and build automation runbooks.

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗