CareerPlanSign in

SRE Engineer

Pune, IND, India💼 Full-time🗓 2026-08-24 → 2026-09-28

Core

Define, monitor, and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets for AI services while maintaining scalable, fault-tolerant infrastructure for AI model serving, training pipelines, and data workflows.

Role type

Senior Site Reliability Engineer (AI/MLOps)

Builds

Production-grade AI microservices, ML workflows, and data pipelines on Azure

Domain

Industrial software, Artificial Intelligence, Cloud Infrastructure

Deliverable

production ML models

Required skills

SLO/SLI/Error Budget management, Incident response and RCA, Azure cloud services, Kubernetes, Docker, Terraform, CI/CD pipelines, Python, Bash, Go, Observability stacks (Prometheus, Grafana, Datadog, Open Telemetry), MLOps practices

Preferred skills

Supervising and logging tools (ELK Stack), Zero-trust architecture principles, Industry compliance standards (SOC 2, ISO 27001, GDPR)

Technologies

Azure, AKS, Azure DevOps, ARM, Terraform, Ansible, Jenkins, GitHub Actions, Prometheus, Grafana, Datadog, Open Telemetry, Docker, Kubernetes, Python, Bash, Go

Responsibilities

Enforce SLOs/SLIs/Error Budgets for AI services; Lead incident response and root cause analysis; Proactively identify and resolve reliability risks; Maintain scalable infrastructure for data pipelines and ML workflows; Handle cloud infrastructure using IaC tools; Oversee Kubernetes clusters and containerized workloads; Maintain observability stacks; Develop intelligent alerting systems; Automate operational tasks via scripting; Build and maintain CI/CD pipelines; Implement MLOps practices; Ensure alignment with security guidelines; Partner with data engineers and scientists; Foster reliability-first culture; Contribute to on-call rotations

Seniority

Senior, hands-on IC

Sourced via siemens_eda · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.