IN-Associate_SRE Engineer_Digital Engineering Transformation_Advisory_Bangalore
Core
Define and implement SLI/SLO frameworks for industrial software platforms, build unified observability stacks across multi-cloud environments, and establish incident management discipline for mission-critical systems.
Role type
Associate Site Reliability Engineer (SRE)
Builds
Observability pipelines, automated runbooks, self-healing patterns, and SRE capability within client organizations
Domain
Industrial software, energy, and OT/IT environments
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
SLI/SLO design, error budget management, observability toolchain (Prometheus, Grafana, Loki, Jaeger, OpenTelemetry), incident management, postmortem facilitation, automation scripting (Python, Go, Bash), GitOps, CI/CD SRE gate experience
Preferred skills
Industrial or OT/IT environment familiarity (energy, manufacturing, connected products)
Technologies
AWS, Azure, OpenTelemetry, Prometheus, Grafana, Loki, Jaeger, ArgoCD, GitHub Actions, Flux
Responsibilities
Define and implement SLI/SLO frameworks for industrial software platforms; Design and deploy observability pipelines; Establish incident management processes and lead blameless postmortems; Build and execute chaos engineering scenarios; Develop automated runbooks and self-healing patterns; Build SRE capability within client organizations
Seniority
Associate, early member of practice