An application made for this job — a tailored resume and cover letter that speak straight to the posting.
The AI Studio is seeking a deeply technical Director of Reinforcement Learning & Agentic Post-Training to lead how our LLM-based agents operate within autonomous supply chain software.
Based at the center of our Model Training Factory, you will guide RL environments, reward models, evaluation harnesses, and production deployment, mentoring senior engineers while balancing performance, cost, and safety.
The AI Studio's mission is to find the fastest possible path to an autonomous supply chain.
We build AI agents, learning systems, model training pipelines, evaluations, simulations, and decision-making systems for some of the hardest problems in global supply chain. The work spans LLMs, reinforcement learning, agentic workflows, software automation, optimization, and production engineering.
In short, we are having a lot of fun.
We are looking for a deeply technical Director of Reinforcement Learning & Agentic Post-Training to lead how Blue Yonder trains LLM-based agents to operate supply chain software.
This role sits at the center of our Model Training Factory, built with NVIDIA, where we develop specialized AI agents for the autonomous supply chain. These agents must reason over supply chain state, use tools, interact with Blue Yonder workflows, execute multi-step operational tasks, and improve through feedback, evaluation, and reinforcement learning.
Tool use is not a side feature here. Our agents must learn to work inside real enterprise software: querying state, proposing actions, invoking APIs, respecting constraints, handling exceptions, escalating uncertainty, and collaborating with human operators. The challenge is not simply making a model sound knowledgeable about supply chain. The challenge is training models that can reliably act.
We are looking for someone who has personally gone through the hard parts: post-training LLMs, designing tool-use environments, building reward models or verifiers, creating evaluations that catch real failures, shipping reinforced models into production, and leading strong machine learning engineers through that process.
This is not a pure research management role, and it is not a project management role. You should be comfortable setting strategy, writing and reviewing technical designs, mentoring senior engineers, challenging weak assumptions, and staying close enough to the work to know whether the system is actually learning.
We want to talk if you:
Supply chains are full of hard AI problems: partial observability, long-horizon consequences, competing objectives, brittle constraints, noisy feedback, and decisions that matter in the real world.
We are not applying reinforcement learning to toy environments. We are training production LLM agents that operate supply chain software through tools, feedback, verification, and reinforcement. The work sits at the intersection of LLMs, agents, reinforcement learning, evaluation, simulation, optimisation, and production engineering.
If you want to build learning systems that leave the lab and operate in one of the world's most complex real-world domains, this is the role.
We want to know the heart of a company, take a look at its values. Ours unite us. They drive our success and the success of our customers.
All qualified applicants will receive consideration for employment without regard to race, colour, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status.
If you want to know the heart of a company, take a look at their values. Ours unite us. They are what drive our success - and the success of our customers. Does your heart beat like ours? Find out here: Core Values
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status.