CareerPlanSign in

Senior Site Reliability Engineer, Production Engineering

India🌐 Remote💼 Full-time🗓 2026-09-21 → 2026-09-25

Core

Senior SRE responsible for maintaining high availability, reliability, and security of large-scale cloud and infrastructure services, specifically focusing on production Kubernetes environments.

Role type

Senior Site Reliability Engineer (Production Engineering)

Builds

Large-scale Kubernetes clusters, high-performance computing environments, and global 24/7 reliability operations

Domain

Cloud infrastructure, High-Performance Computing (HPC), Kubernetes, Bare-metal systems

Deliverable

production ML models | infrastructure

Required skills

Kubernetes administration, Linux systems administration, Networking (DNS, DHCP, IP tables, routing, firewalls), Incident management, Observability, Automation/Scripting, CI/CD pipelines, Troubleshooting complex infrastructure issues

Preferred skills

GPU/DPU hardware knowledge, Python/Golang/Rust programming, Architecture of large-scale Kubernetes environments, SLURM experience

Technologies

Kubernetes, SLURM, Jenkins, ArgoCD, Linux, Python, Golang, Rust

Responsibilities

Administer and maintain large-scale Kubernetes clusters and infrastructure; Automate operational processes to reduce manual tasks; Proactively detect and respond to production incidents using monitoring and observability; Lead incident management calls and coordinate resolution of critical issues; Analyze logs and metrics to troubleshoot complex problems; Contribute to the architecture and deployment of large-scale Kubernetes environments

Seniority

Senior, hands-on IC

Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.