Director of AI Workload Intelligence
Core
Lead global initiatives to analyze LLM and multimodal workloads for HBF (High Bandwidth Fabric) applicability, defining data placement criteria and guiding product strategy.
Role type
Director of AI Workload Intelligence
Builds
HBF-aware memory solutions for AI data centers
Domain
Semiconductor / AI Infrastructure
Deliverable
production ML models
Required skills
AI system architecture leadership, LLM serving, AI model analysis, heterogeneous accelerator software, Transformers, MoE, embedding/retrieval workloads, latency/throughput/memory footprint analysis
Preferred skills
Compiler/Runtime and AI Inference SW Stack integration
Technologies
PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, ONNX Runtime, Triton Inference Server
Responsibilities
Lead global collaboration and drive AI workload-based HBF technology development from research to productization; Analyze AI models, algorithms, and workload trends for HBF applicability; Classify HBF-relevant memory objects such as weights, KV cache, activations, embeddings, adapters, and MoE experts; Define model-level criteria for data placement, caching, prefetching, and offload decisions; Analyze HBF integration points in vLLM, SGLang, TensorRT-LLM, PyTorch, and ONNX Runtime; Identify HBF-applicable AI model families and inference data objects; Develop memory-use taxonomy and HBF applicability criteria; Provide model-driven inputs to Compiler/Runtime and AI Inference SW Stack teams
Seniority
Director, global strategy & execution