Director, Reinforcement Learning & Agentic Post-Training
Core
Lead the technical strategy and execution for training LLM-based agents to operate autonomous supply chain software using reinforcement learning and post-training techniques.
Role type
Director, Reinforcement Learning & Agentic Post-Training
Builds
Production LLM agents that reason over supply chain state, use tools, interact with enterprise workflows, and execute multi-step operational tasks.
Domain
Supply Chain Management / Artificial Intelligence / Reinforcement Learning
Deliverable
production ML models
Required skills
Reinforcement Learning (PPO, GRPO, offline RL), LLM post-training (SFT, DPO, RLHF/RLAIF), Tool-use environment design, Reward modeling and verifiers, Evaluation frameworks for agent behavior, Python, PyTorch, Team leadership for ML engineers, Production engineering constraints (latency, cost, safety)
Preferred skills
NVIDIA stack (Nemotron, NeMo, Megatron, vLLM, Ray), Distributed training, Large-scale inference systems, Simulated enterprise software environments, Supply chain domain knowledge (planning, warehouse, transportation), Agent safety systems design
Technologies
Python, PyTorch, NVIDIA Nemotron, NVIDIA NeMo, Megatron, vLLM, Ray
Responsibilities
Lead technical strategy for RL and post-training; Build and manage ML engineering teams; Design training environments for agent tool use; Develop reward models and evaluation harnesses; Define operational quality metrics for agents; Partner with domain experts to create trainable workflows; Guide model improvement across optimization techniques; Establish engineering standards for reproducibility and rollout safety.
Seniority
Director, hands-on technical leadership