CareerPlanGet AI match score →

Lead Site Reliability Engineer

Bangalore, IND💼 Full-time🗓 2026-07-15 → 2026-07-31

Core

Lead Site Reliability Engineer responsible for ensuring 99.9%+ uptime, managing SLOs/SLIs, and driving reliability for critical services.

Role type

Lead Site Reliability Engineer

Builds

resilient, scalable, high-performance services and infrastructure

Domain

Cloud infrastructure, distributed systems, marketing technology

Deliverable

production ML models | infrastructure

Required skills

SLO/SLI management, incident response, root cause analysis, capacity planning, observability design, chaos engineering, Infrastructure as Code, Linux systems administration, cloud platform expertise, container orchestration, programming in Python or Go, shell scripting, distributed tracing integration

Preferred skills

distributed systems experience, statistical analysis of metrics, high-performance low-latency systems knowledge, on-call rotation experience, production-level code writing, Chaos Engineering drill execution

Technologies

AWS, Kubernetes, EKS, Fargate, OpenTelemetry, Honeycomb, Grafana, Prometheus, Thanos, ELK, Loki, Terraform, Pulumi, Chaos Mesh, AWS Fault Injection Simulator

Responsibilities

Implement and manage SLOs, SLIs, and error budgets; Develop resilient systems ensuring 99.9%+ uptime; Lead incident response and post-incident reviews; Automate incident detection and response; Write software to support reliability needs; Design and implement full observability; Perform capacity planning and performance testing; Collaborate on building reliable services; Ensure best practices in infrastructure design and deployment; Champion Infrastructure as Code; Participate in chaos engineering initiatives; Participate in on-call rotation; Drive advanced alerting and anomaly detection

Seniority

Senior, hands-on IC with leadership responsibilities

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗