Site Reliability Engineer (w/m/d)
Core
Building and maintaining the infrastructure for Managed Nextcloud, Nextcloud Workspace, IONOS GPT, and other web services on a Kubernetes platform.
Role type
Site Reliability Engineer (IC)
Builds
Production-grade containerized web services and platform infrastructure
Domain
Cloud infrastructure, Kubernetes, SaaS
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes, Linux, Infrastructure as Code (Terraform), CI/CD pipelines, scripting (Go/Python/Bash), distributed systems debugging, monitoring and logging (Prometheus, Grafana, ELK)
Preferred skills
Helm Charts, building Operators, experience with high-availability environments
Technologies
Kubernetes, Terraform, GitLab CI/CD, ArgoCD, Prometheus, Grafana, ELK-Stack, Go, Python, Bash
Responsibilities
Develop and integrate new products/services into the Kubernetes and cloud infrastructure; ensure stable and secure operation of the product platform; automate infrastructure provisioning and management; analyze and resolve complex issues in distributed systems; develop and maintain monitoring, logging, and alerting solutions.
