CareerPlanSign in

HPC/ML Infrastructure Engineer

San Francisco or Tokyo💼 Full-time🗓 2026-06-15 → 2026-09-26

Core

Lead bringup, administration, and operations for a large-scale GPU cluster dedicated to training anime AI models.

Role type

Senior HPC/ML Infrastructure Engineer

Builds

Large-scale GPU clusters for anime AI model training

Domain

High-Performance Computing (HPC) / Artificial Intelligence

Deliverable

infrastructure

Required skills

SLURM management, Kubernetes (K8s), Linux system administration, storage systems (WEKA, VAST, Ceph), network configuration, Ansible, Terraform, Grafana, Prometheus

Preferred skills

Warewulf, MAAS, Tailscale, LDAP, physical server racking and stacking

Technologies

SLURM, Kubernetes, Ansible, WEKA, VAST, Ceph, Tailscale, Grafana, Prometheus, Warewulf, MAAS

Responsibilities

Deploy and manage SLURM on K8s clusters, provision infrastructure using Warewulf/MAAS/Ansible, manage parallel filesystems and network transmission, troubleshoot hardware and OS issues, rack and stack physical GPU nodes

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.