CareerPlanSign in

DevOps Engineer

Singapore💼 Full-time🗓 2026-05-21 → 2026-08-06

Core

Design, deploy, and support large-scale distributed GPU clusters for AI and ML workloads, managing automation and monitoring for both on-prem and cloud environments.

Role type

DevOps Engineer (GPU-as-a-Service)

Builds

GPU clusters and infrastructure for AI/HPC workloads

Domain

Cloud computing, AI/ML infrastructure, High-Performance Computing (HPC)

Deliverable

production ML models | infrastructure

Required skills

Linux system administration, Kubernetes, Terraform, Ansible, Jenkins, Python, Bash, CI/CD, GPU architecture, NVIDIA GPUs, Zabbix, Prometheus, Slurm, IB networking, OS optimization, CUDA

Preferred skills

TensorFlow, Pytorch, security best practices for multi-tenant environments

Technologies

Ubuntu, CentOS, Rocky Linux, Jenkins, Kubernetes, Ansible, Terraform, Zabbix, Prometheus, NVIDIA DCGM, Slurm, CUDA, IB networking, TensorFlow, PyTorch

Responsibilities

Design and deploy large-scale distributed GPU clusters; Manage and automate provisioning of GPU resources; Design and implement CI/CD pipelines for AI models; Monitor cluster usage, health, and performance; Optimize system parameters for AI workload performance; Troubleshoot compute resource system-level issues; Implement security best practices for multi-tenant GPUaaS environments; Provide technical support and guidance to users of GPU-accelerated systems.

Seniority

Mid-level, hands-on IC

Sourced via adzuna · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.