DevOps Engineer, GPUaaS
Core
Implement processes and integration of operations to advance customer's AI and HPC capabilities within a GPU-as-a-Service (GPUaaS) environment.
Role type
DevOps Engineer (GPUaaS)
Builds
Large-scale, distributed GPU clusters for AI and ML workloads
Domain
AI/HPC cloud platforms, GPU infrastructure
Deliverable
production ML models | infrastructure
Required skills
Linux system administration, DevOps tools (Jenkins, Kubernetes, Ansible, Terraform), CI/CD pipelines, scripting (Python, Bash), monitoring (Zabbix, Prometheus), GPU architecture and NVIDIA GPUs, cloud architectures (IaaS, PaaS)
Preferred skills
Collective communications (MPI, RDMA, NCCL), HPC workload managers (Slurm), AI & HPC networking (InfiniBand, RoCE, DPUs), AI frameworks (TensorFlow, PyTorch), Docker/containers
Technologies
Ubuntu/CentOS/Rocky Linux, Jenkins, Kubernetes, Ansible, Terraform, Python, Bash, Zabbix, Prometheus, Slurm, InfiniBand, RoCE, DPUs, TensorFlow, PyTorch, NVIDIA DCGM
Responsibilities
Design, deploy and support large-scale, distributed GPU clusters; Manage and automate provisioning of GPU resources in on-prem and cloud platforms; Design, implement and manage CI/CD pipelines for AI models and GPU-accelerated applications; Monitor cluster usage, health, performance and availability; Troubleshoot compute resource system level issues such as Slurm, Kubernetes, GPU drivers, CUDA, IB networking; Optimize system parameters for AI workload performance; Conduct GPU cluster benchmark and keep up with the latest advancements in GPU technology; Set up monitoring and logging for GPU resources; Implement security best-practices for multi-tenant GPU-as-a-Service environment; Provide technical support and guidance to users of GPU-accelerated systems; Collaborate with software and administrators to streamline workflows; Work with senior DevOps engineers to identify bottlenecks and improve development and operational processes
Seniority
Mid-level to Senior, hands-on IC