(Senior) Site Reliability Engineer / Distributed Cloud - STACKIT (m/w/d)
Core
Operate and optimize complex distributed cloud platforms (Kubernetes, KubeVirt, Cilium, Ceph, Talos) ensuring end-to-end stability, scalability, and cost efficiency for retail and external clients.
Role type
Senior Site Reliability Engineer (Distributed Cloud)
Builds
High-availability cloud infrastructure and monitoring systems for Lidl, Kaufland, Schwarz Produktion, PreZero, and external European enterprises.
Domain
Retail technology / Distributed Cloud Systems
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Cloud infrastructure operations, Golang, Synthetic monitoring, SLO definition, Toil reduction via automation
Preferred skills
Ceph, Cilium, KubeVirt, Talos, Systems programming
Technologies
Kubernetes, KubeVirt, Cilium, Ceph, Talos, Golang
Responsibilities
Operate and optimize complex distributed cloud platforms; Develop and maintain monitoring and logging systems; Implement synthetic monitoring and trace tests; Define and monitor Service Level Objectives (SLOs); Automate toil reduction through code.
Seniority
Senior, hands-on IC
