Site Reliability Engineer / SRE - Cloud Storage - STACKIT (m/w/d)
Core
Maintain and optimize the stability and availability of high-availability, resilient storage infrastructure (Block, Object, Backup, and File Storage) for internal retail operations and external clients.
Role type
Senior Site Reliability Engineer (Cloud Storage)
Builds
Robust and efficient storage architecture and services for retail partners (Lidl, Kaufland) and external European enterprises.
Domain
Cloud Storage Infrastructure & Retail IT
Deliverable
production ML models | product features | infrastructure
Required skills
Storage product expertise (NetApp, Cohesity, Pure, Ceph), Cloud environment architecture, Infrastructure automation (Golang, Python, Bash, Ansible), Container orchestration (k8s), Monitoring & Alerting (Prometheus, Grafana, Elasticsearch), API development (REST)
Preferred skills
Performance analysis, Capacity planning, Incident response, Troubleshooting
Technologies
Golang, Python, Bash, Ansible, Kubernetes, Prometheus, Grafana, Elasticsearch, NetApp, Cohesity, Pure, Ceph
Responsibilities
Proactive monitoring and resolution of storage infrastructure incidents, Automating deployment and operational processes, Designing and maintaining robust storage architecture, Analyzing performance and planning capacity for future scaling, Conducting incident and postmortem analysis
Seniority
Senior, hands-on IC