Site Reliability Engineer
Core
Lead incident response efforts, ensure system reliability and scalability, and implement monitoring solutions for P&G's warehousing IT operations.
Role type
Site Reliability Engineer (Incident Response)
Builds
Highly available and efficient infrastructure for warehousing systems
Domain
Retail / Warehousing IT Operations
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux/Unix administration, cloud platforms (AWS/Azure/GCP), SAP, Terraform, Python/C#, networking protocols, Docker, Kubernetes, SQL, Prometheus, Grafana, root cause analysis
Preferred skills
Warehousing Management Systems (RTCIS, PrIME)
Responsibilities
Lead incident response and resolution, conduct root cause analysis, implement monitoring solutions, optimize system architecture, collaborate with cross-functional teams
Seniority
Mid-level, hands-on IC