CareerPlanGet AI match score →

Engineering Manager, Site Reliability Engineering (SRE)

Chennai India💼 Full-time🗓 2026-07-09 → 2026-07-31

Core

Lead a team of Site Reliability and Infrastructure Engineers to drive reliability, observability, automation, and operational readiness for SaaS infrastructure supporting healthcare service operations.

Role type

Engineering Manager, Site Reliability Engineering (SRE)

Builds

Highly available SaaS infrastructure, operational tooling, observability solutions, and automation capabilities

Domain

Healthcare / Cloud Infrastructure / Linux Systems

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Linux system administration, Infrastructure-as-Code (Terraform, Puppet, Ansible), observability platforms (metrics, logs, traces), incident management, team leadership, Agile delivery, Python/Go/Bash scripting, Kubernetes/EKS operations

Preferred skills

None explicitly stated

Technologies

Puppet, Ansible, Terraform, New Relic, Prometheus, Alertmanager, OpenSearch, Grafana, Icinga, Unified Assurance, Kubernetes, Amazon EKS

Responsibilities

Lead, coach, and mentor a team of SREs; review designs and troubleshoot complex issues; drive observability strategy and alert quality; manage provisioning and lifecycle of Linux systems; partner with engineering teams on SaaS/hybrid cloud environments; reduce operational toil through automation; participate in incident response and post-incident analysis; develop operational metrics and documentation

Seniority

Manager, hands-on technical leadership

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗