HPC/ML Infrastructure Engineer
Core
Lead bringup, administration, and operations for a large-scale GPU cluster dedicated to training anime AI models.
Role type
Senior HPC/ML Infrastructure Engineer
Builds
Large-scale GPU clusters for anime AI model training
Domain
High-Performance Computing (HPC) / Artificial Intelligence
Deliverable
infrastructure
Required skills
SLURM management, Kubernetes (K8s), Linux system administration, storage systems (WEKA, VAST, Ceph), network configuration, Ansible, Terraform, Grafana, Prometheus
Preferred skills
Warewulf, MAAS, Tailscale, LDAP, physical server racking and stacking
Technologies
SLURM, Kubernetes, Ansible, WEKA, VAST, Ceph, Tailscale, Grafana, Prometheus, Warewulf, MAAS
Responsibilities
Deploy and manage SLURM on K8s clusters, provision infrastructure using Warewulf/MAAS/Ansible, manage parallel filesystems and network transmission, troubleshoot hardware and OS issues, rack and stack physical GPU nodes
Seniority
Senior, hands-on IC