CareerPlanSign in

Senior Site Reliability Engineer

Bengaluru💼 Full-time🗓 2025-09-25 → 2026-09-25

Core

Own reliability across a cloud-native and AI-driven platform, building systems that scale and self-heal using LLM-powered automation.

Role type

Senior Site Reliability Engineer (AI/LLM focus)

Builds

Cloud-native infrastructure, AI agents for DevOps/SRE workflows, and LLM-powered operational tooling

Domain

Cloud infrastructure, AI/LLM, Kubernetes operations

Deliverable

production ML models | infrastructure

Required skills

AWS infrastructure at scale, Kubernetes (production-grade), distributed systems debugging, Python or Go, monitoring and alerting systems

Preferred skills

LLM APIs (OpenAI), agent frameworks (LangChain, AutoGen), AI agents for DevOps, RAG systems, vector databases, AIOps

Technologies

AWS, Kubernetes, OpenAI API, LangChain, AutoGen, Pinecone, Weaviate, Prometheus, Grafana, ELK, OpenSearch, OpenTelemetry, Terraform

Responsibilities

Own uptime and performance of services on AWS + Kubernetes, design self-healing infrastructure, build LLM-powered tooling for alert triage and incident analysis, manage and scale Kubernetes workloads, build observability systems, define SLOs/SLAs, automate infrastructure, lead incident response, introduce chaos engineering

Seniority

Senior, hands-on IC

Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.