Site Reliability Engineer - Warehousing IT Operations
Core
Lead incident response efforts and ensure high system availability for warehousing IT operations, collaborating with software engineers and DevOps teams to maintain stability and uptime.
Role type
Managerial Site Reliability Engineer (Warehousing IT Operations)
Builds
Robust monitoring tools, automated incident response systems, and resilient system architectures for warehousing services.
Domain
Warehousing IT Operations / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Incident response leadership, root cause analysis, Linux/Unix administration, cloud platforms (AWS/Azure/GCP), infrastructure-as-code (Terraform), scripting (Python/C#), networking protocols, containerization (Docker/Kubernetes), SQL, monitoring tools (Prometheus/Grafana), security best practices.
Preferred skills
Warehousing Management Systems (RTCIS/PrIME) experience, Warehousing Operations background.
Responsibilities
Lead incident response and resolution, conduct root cause analysis, optimize system architecture for scalability, implement comprehensive monitoring solutions, collaborate cross-functionally on resilient system design, mentor team members on SRE best practices.
Seniority
Managerial, hands-on IC with team leadership