DevOps Engineer
Core
Design, deploy, and support large-scale distributed GPU clusters for AI and ML workloads, managing automation and monitoring for both on-prem and cloud environments.
Role type
DevOps Engineer (GPU-as-a-Service)
Builds
GPU clusters and infrastructure for AI/HPC workloads
Domain
Cloud computing, AI/ML infrastructure, High-Performance Computing (HPC)
Deliverable
production ML models | infrastructure
Required skills
Linux system administration, Kubernetes, Terraform, Ansible, Jenkins, Python, Bash, CI/CD, GPU architecture, NVIDIA GPUs, Zabbix, Prometheus, Slurm, IB networking, OS optimization, CUDA
Preferred skills
TensorFlow, Pytorch, security best practices for multi-tenant environments
Technologies
Ubuntu, CentOS, Rocky Linux, Jenkins, Kubernetes, Ansible, Terraform, Zabbix, Prometheus, NVIDIA DCGM, Slurm, CUDA, IB networking, TensorFlow, PyTorch
Responsibilities
Design and deploy large-scale distributed GPU clusters; Manage and automate provisioning of GPU resources; Design and implement CI/CD pipelines for AI models; Monitor cluster usage, health, and performance; Optimize system parameters for AI workload performance; Troubleshoot compute resource system-level issues; Implement security best practices for multi-tenant GPUaaS environments; Provide technical support and guidance to users of GPU-accelerated systems.
Seniority
Mid-level, hands-on IC