CareerPlanGet AI match score →

Manager, Site Reliability Engineering

Noida💼 Full-time🗓 2026-07-01 → 2026-07-31

Core

Lead the reliability, stability, and operational excellence of enterprise platforms, owning 24x7 incident management and SRE engineering efforts for high system availability.

Role type

Senior IC SRE Manager

Builds

Production systems, AI/ML workloads, and automated operational solutions for global brands

Domain

Cloud Infrastructure & Site Reliability Engineering

Deliverable

production ML models | infrastructure

Required skills

SRE leadership, incident management, root cause analysis, observability, automation, AWS, Kubernetes, microservices, release management, SLO/SLI design, cloud cost optimization

Preferred skills

AI/ML workload support, security best practices, team mentoring, strategic alignment

Technologies

AWS, Kubernetes (EKS), Datadog, CloudWatch, ELK, Prometheus, Grafana

Responsibilities

Own end-to-end reliability of production systems within SLAs; Lead and govern a 24x7x365 incident management team; Drive blameless RCA culture and action item closure; Improve observability and reduce alert noise; Drive automation to reduce operational toil; Mentor a team of ~14 engineers; Optimize cloud usage and reduce spend.

Seniority

Senior, hands-on IC with people leadership

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗