Site Reliability Engineer – Bygg Norges Sovereign AI Cloud
Core
Build and maintain a Norwegian Sovereign AI Cloud platform ensuring stable operations and continuous improvement for advanced AI workloads.
Role type
Senior Site Reliability Engineer (SRE)
Builds
GPU-activated, cloud-based solutions for advanced AI/ML workloads
Domain
Sovereign AI Cloud, High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Observability (metrics, logging, tracing), Infrastructure as Code (IaC), GitOps, CI/CD, Capacity planning, Disaster recovery, Backup management, AI/ML platform experience
Preferred skills
OpenShift storage (CSI, Ceph, NFS, S3), Prometheus, Grafana, OpenTelemetry, Ansible, High-performance computing experience
Technologies
Kubernetes, OpenShift, Grafana, Prometheus, OpenTelemetry, Ansible, GitOps, CI/CD, Ceph, NFS, S3
Responsibilities
Implement and operate observability solutions in Kubernetes, Monitor and handle events using monitoring and logging tools, Automate operations and configuration via IaC and GitOps, Administer storage, backup, and recovery solutions, Plan and execute capacity and scaling activities, Analyze performance and usage data for platform and AI workloads
Seniority
Senior, hands-on IC
