CareerPlanSign in

Director of AI Workload Intelligence

San Jose, CA💼 Full-time💰 $250,000–$250,000🗓 2026-09-01 → 2026-09-26

Core

Lead global initiatives to analyze LLM and multimodal workloads for HBF (High Bandwidth Fabric) applicability, defining data placement criteria and guiding product strategy.

Role type

Director of AI Workload Intelligence

Builds

HBF-aware memory solutions for AI data centers

Domain

Semiconductor / AI Infrastructure

Deliverable

production ML models

Required skills

AI system architecture leadership, LLM serving, AI model analysis, heterogeneous accelerator software, Transformers, MoE, embedding/retrieval workloads, latency/throughput/memory footprint analysis

Preferred skills

Compiler/Runtime and AI Inference SW Stack integration

Technologies

PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, ONNX Runtime, Triton Inference Server

Responsibilities

Lead global collaboration and drive AI workload-based HBF technology development from research to productization; Analyze AI models, algorithms, and workload trends for HBF applicability; Classify HBF-relevant memory objects such as weights, KV cache, activations, embeddings, adapters, and MoE experts; Define model-level criteria for data placement, caching, prefetching, and offload decisions; Analyze HBF integration points in vLLM, SGLang, TensorRT-LLM, PyTorch, and ONNX Runtime; Identify HBF-applicable AI model families and inference data objects; Develop memory-use taxonomy and HBF applicability criteria; Provide model-driven inputs to Compiler/Runtime and AI Inference SW Stack teams

Seniority

Director, global strategy & execution

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.