Director, Reinforcement Learning & Agentic Post-Training

Blue Yonder

Paris

Sur place

EUR 90 000 - 120 000

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Résumé du poste

Blue Yonder is looking for a Director of Reinforcement Learning & Agentic Post‑Training in Paris to lead the training of LLM-based agents for autonomous supply chains. The role includes managing a team of machine-learning engineers and establishing technical strategies for agent development.

Ideal candidates should have experience leading high-performing teams in machine learning and reinforcement learning, and proficiency in Python and PyTorch. Join us to tackle real-world AI challenges in global supply chain management.

Qualifications

  • Experience leading teams in deploying LLM models with reinforcement learning.
  • Strong background in machine learning engineering.
  • Ability to reason about production constraints.

Responsabilités

  • Lead technical strategy for tool-using LLM agents.
  • Manage team focused on agent training and evaluation systems.
  • Partner with experts and document learning outcomes.

Connaissances

Reinforcement Learning
Machine Learning Engineering
Python
PyTorch
Team Leadership

Outils

NVIDIA
APIs

Description du poste

About The AI Studio

The AI Studio's mission is to find the fastest possible path to an autonomous supply chain. We build AI agents, learning systems, model training pipelines, evaluations, simulations, and decision‑making systems for some of the hardest problems in global supply chain. The work spans LLMs, reinforcement learning, agentic workflows, software automation, optimization, and production engineering. In short, we are having a lot of fun.

Your Mission

We are seeking a deeply technical Director of Reinforcement Learning & Agentic Post‑Training to lead how Blue Yonder trains LLM‑based agents to operate supply chain software. This role sits at the center of our Model Training Factory, built with NVIDIA, where we develop specialized AI agents for the autonomous supply chain. These agents must reason over supply chain state, use tools, interact with Blue Yonder workflows, execute multi‑step operational tasks, and improve through feedback, evaluation, and reinforcement learning. Tool use is not a side feature here. Our agents must learn to work inside real enterprise software: querying state, proposing actions, invoking APIs, respecting constraints, handling exceptions, escalating uncertainty, and collaborating with human operators. The challenge is training models that can reliably act.

What You'll Do
  • Lead the technical strategy for reinforcement learning, post‑training, and tool‑using LLM agents within the AI Studio.
  • Build and manage a team of machine‑learning engineers working on agent training, RL environments, reward modeling, evaluation, data generation, and training infrastructure.
  • Design environments where LLM agents learn to operate Blue Yonder software through APIs, tools, workflows, simulations, and human feedback.
  • Develop training and evaluation systems for multi‑step supply chain workflows across planning, warehouse management, transportation, commerce, and network operations.
  • Define what "good" looks like for operational agents: correct tool use, constraint adherence, business outcome quality, latency, cost, robustness, escalation behavior, and human trust.
  • Build reward models, verifiers, preference pipelines, automated graders, and evaluation harnesses for agent behavior.
  • Create evaluation frameworks that measure real agent performance, including tool‑call correctness, workflow completion, recovery from bad state, long‑horizon reliability, and failure modes.
  • Partner with product, engineering, architecture, and domain experts to turn real supply chain workflows into trainable agent environments.
  • Guide model improvement across supervised fine‑tuning, preference optimization, reinforcement learning from human or AI feedback, rejection sampling, synthetic data generation, and policy optimisation.
  • Make practical technical tradeoffs between model capability, inference cost, latency, reliability, product timelines, and operational safety.
  • Establish engineering standards for experiment tracking, reproducibility, observability, rollout safety, and production monitoring.
  • Document what works and what fails so the team compounds learning over time.
What We're Looking For
  • Have led a team to ship LLM models trained with reinforcement learning, SFT, DPO, RLHF/RLAIF, and other post‑trained models in production.
  • Have led a team to train models to use tools, call APIs, interact with software environments, or complete multi‑step tasks.
  • Have a strong machine‑learning engineering background and can credibly lead engineers because you have built systems like this yourself.
  • Have managed or technically led high‑performing reinforcement learning ML engineering teams.
  • Are highly proficient in Python and PyTorch.
  • Understand modern LLM post‑training workflows, including supervised fine‑tuning, preference data, reward modeling, policy optimisation, evaluation, and deployment.
  • Have hands‑on experience with reinforcement learning methods such as reward shaping, PPO‑style optimisation, GRPO, offline RL, policy evaluation, rejection sampling, or environment design.
  • Know how to evaluate open‑ended agent behaviour beyond static benchmark scores.
  • Can reason about production constraints: latency, inference cost, safety, observability, rollback, and reliability.
  • Can balance frontier‑oriented exploration with shipping production systems.
  • Are comfortable with ambiguity but intolerant of unsound technical thinking.
  • Care about engineering craft, reproducibility, and learning velocity.
  • Are curious about why systems work, not just whether a metric moved.
Bonus Points
  • Experience building simulated or sandboxed enterprise software environments for agent training.
  • Experience with NVIDIA Nemotron, NVIDIA NeMo, Megatron, vLLM, Ray, distributed training, or large‑scale inference systems.
  • Experience with warehouse management, supply chain planning, transportation, merchandising, logistics, operations research, or enterprise workflow automation.
  • Experience designing agent safety systems, including permissioning, action validation, approval flows, uncertainty escalation, and audit trails.
  • Evidence of technical taste through papers, open‑source contributions, internal platforms, side projects, or shipped systems that show deep curiosity about model behaviour.
What Makes This Role Different

Supply chains are full of hard AI problems: partial observability, long‑horizon consequences, competing objectives, brittle constraints, noisy feedback, and decisions that matter in the real world. We are not applying reinforcement learning to toy environments. We are training production LLM agents that operate supply chain software through tools, feedback, verification, and reinforcement. The work sits at the intersection of LLMs, agents, reinforcement learning, evaluation, simulation, optimisation, and production engineering. If you want to build learning systems that leave the lab and operate in one of the world's most complex real‑world domains, this is the role.

Equal Opportunity

All qualified applicants will receive consideration for employment without regard to race, colour, religion, sex, sexual orientation, gender identity, national origin, disability, or protected veteran status. All legally protected characteristics are respected.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Director, Agentic RL for Autonomous Supply Chains
Director, Agentic RL for Autonomous Supply Chains

Blue Yonder • Paris

Sur place
EUR 90 000 - 120 000
Technical Staff Member – Agent
Technical Staff Member – Agent

Jobtailor • Paris

Sur place
EUR 120 000 - 180 000
ML Infrastructure Engineer
ML Infrastructure Engineer

White Circle • Paris

Hybride
EUR 157 000 - 307 000
Relocation package
Hybrid Paris/London work
Medical insurance
+2
ML Infrastructure Engineer
ML Infrastructure Engineer

White Circle • Paris

Hybride
EUR 70 000 - 90 000
Comprehensive medical insurance
Paid time off
Team off-sites
+1
Applied AI Engineer
Applied AI Engineer

Norbert Health • Paris

Sur place
EUR 70 000 - 90 000
Competitive salary and equity
High autonomy and technical ownership
Transparent, mission-driven culture
Research Engineer — Robot Learning
Research Engineer — Robot Learning

Bleu Robotics • Paris

Sur place
EUR 90 000 - 130 000
Ambitious research mission
Advanced robots fleet
Research that ships to production
+5
AI Engineer
AI Engineer

BeTomorrow – SARL • Bordeaux

Sur place
EUR 60 000 - 80 000
Data & ML Infrastructure Lead
Data & ML Infrastructure Lead

S27a • Paris

Sur place
EUR 120 000 - 160 000
ML Infrastructure Engineer
ML Infrastructure Engineer

Visa Hunt • Paris

Hybride
EUR 85 000 - 135 000
Relocation package
Comprehensive medical insurance (Paris
Hardware and tools provided
+2
Senior Applied AI Researcher — Post-Training (RL) & Evaluation
Senior Applied AI Researcher — Post-Training (RL) & Evaluation

Lumos • Paris

Sur place
EUR 150 000 - 173 000
Performance bonus
Equity options