Director, Model Behavior & Evaluation Systems
Core
Director of Model Behavior & Evaluation Systems owning the behavioral quality system for Blue Yonder's LLM agents operating in supply chain workflows.
Role type
Director, technical leadership for AI agents and model evaluation
Builds
LLM agents, learning systems, model training pipelines, evaluations, simulations, and decision-making systems for global supply chain
Domain
Supply chain management + Large Language Models (LLMs) and AI Agents
Deliverable
production ML models | product features
Required skills
LLM product leadership, model behavior & evaluation system design, LLM tool/function calling, post-training workflows (SFT, RLHF/RLAIF), technical fluency in Python/PyTorch/Hugging Face/OpenAI Agents SDK, red-teaming and behavioral risk management, cross-functional team leadership, product judgment for agent behavior
Preferred skills
Experience scaling evaluation/launch-readiness functions, frontier lab or enterprise AI platform background
Technologies
Python, PyTorch, Hugging Face Transformers, Hugging Face Datasets, NVIDIA NeMo RL, OpenAI Agents SDK, Langfuse, LLM evaluation harnesses
Responsibilities
Own behavioral quality bar for LLM agents in supply chain workflows; build evaluation systems including datasets, graders, rubrics, and launch gates; define launch criteria for operational correctness and tool-use accuracy; convert traces and feedback into behavior specs and training data; partner with RL/post-training teams on SFT and reward modeling; lead red-teaming for hallucinations and unsafe recommendations; manage model release readiness and post-launch monitoring; build and lead the model behavior function.
Seniority
Director, technical leadership with hands-on depth