Identity & Access Management Site Reliability Engineer
Core
Ensure smooth operation of production systems for Digital Products by combining engineering principles, operational discipline, and automation to maintain availability, reliability, and scalability.
Role type
Site Reliability Engineer (SRE)
Builds
Digital Products and their underlying platform infrastructure
Domain
Consumer Goods / IT Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Python, Bash, PowerShell, SQL, Linux/Unix administration, AWS/Azure/GCP, Prometheus, Grafana, ITIL Service Management, SLO/SLI/SLA definition, Incident Management, Root Cause Analysis, Automation scripting, Network protocols, Database management, Software Development Lifecycle (SDLC), Agile/Scrum, DevOps
Preferred skills
IAM concepts (RBAC, ACL), Ping Authentication/Ping Federate/Ping Radius, Centrify, Ansible, Terraform, Infrastructure as Code (IaC), Chaos Engineering, Test Driven Development (TDD), ITIL 4 Foundation certification, SRE Fundamentals certification, DevOps Foundation certification, PSM1 certification
Technologies
Python, Bash, PowerShell, SQL, Linux, Unix, AWS, Azure, GCP, Prometheus, Grafana, Snow, Ansible, Terraform, Ping ID, Ping Federate, Ping Radius, Centrify
Responsibilities
Monitor system alerts and respond to incidents to meet SLAs/SLOs; Perform root cause analysis and create postmortems; Automate repetitive tasks and maintain monitoring/logging systems; Provide expert guidance to stakeholders on supportability and scalability; Administer platform infrastructure and oversee change management; Oversee vendor management and define strategies for digital products; Represent Operations in PoC and Pilots; Drive observability implementations with monitoring, metrics, and knowledge sharing
Seniority
Experienced Professional