VP, Product AI/ ML
Core
Architect the product strategy and engineering execution for the Research Training Stack, enabling frontier AI model pre-training and post-training at massive scale.
Role type
VP, Product AI/ML (Infrastructure & Orchestration)
Builds
High-performance orchestration tools (SUNK/Slurm on Kubernetes), automated training evaluation frameworks, and RL/RLHF pipelines for AI research labs.
Domain
AI Infrastructure / High-Performance Computing (HPC) / Cloud Native
Deliverable
production ML models
Required skills
Engineering leadership, Slurm, Kubernetes, distributed training cluster networking (InfiniBand/RDMA), frontier model research lifecycle, multi-thousand GPU cluster scaling, strategic product vision
Preferred skills
Experience with NVIDIA H100/Blackwell/Rubin architectures, RL/RLHF pipeline design
Technologies
Slurm, Kubernetes, InfiniBand, RDMA, NVIDIA GPUs
Responsibilities
Oversee the evolution of SUNK (Slurm on Kubernetes) for deterministic bare-metal performance; drive development of next-generation orchestrators and automated training evaluation frameworks; build infrastructure for Reinforcement Learning (RL) and RLHF pipelines; act as primary technical partner for lead researchers at global AI labs.
Seniority
Executive (VP), Strategy & Hands-on IC