Machine Learning Engineer - Reinforcement Learning

JDA Software

Paris

Sur place

EUR 70 000 - 110 000

Plein temps

14 jours+
Générateur de candidature

N’envoyez pas un CV générique — générez un CV et une lettre de motivation adaptés à ce poste précis.

Passez les filtres ATS

Résumé du poste

Blue Yonder is seeking an ambitious ML Engineer in Paris to help build training, evaluation, and tooling systems for AI-powered decision-making in supply chain. The role centers on LLMs, agent environments, and reinforcement learning, with emphasis on robust evaluation and data pipelines.

You will design experiments, ship production code, and shape how LLMs are used inside agent-based systems. A strong background in RL training with LLMs, reward modeling, and Python/PyTorch is required, with a

Qualifications

  • You've trained or fine-tuned LLMs.
  • Experience with AI agents and RL environments in production.
  • Proficiency in Python and PyTorch.
  • Hands-on experience with RL techniques (reward shaping, policy optimization).
  • Ability to balance research exploration with shipping working code.

Responsabilités

  • Design and implement LLM-powered agent environments for supply chain decision-making.
  • Fine-tune and evaluate LLMs for domain-specific reasoning and decision support.
  • Design, test, and iterate on reward functions that capture the behaviors we want from LLM agents.
  • Review LLM traces and rollouts to understand model reasoning, failure modes, reward hacking, and shortcut behaviour.
  • Identify when an LLM is exploiting the reward function or optimizing for proxy metrics.
  • Improve reward models, environment design, prompts, tools, and feedback loops.
  • Build evaluation frameworks to measure model quality, agent performance, robustness, and failure modes.
  • Create data pipelines for training, fine-tuning, preference data, synthetic data generation, and human feedback collection.
  • Develop tooling that improves how the team builds, tests, debugs, and deploys AI-assisted workflows.
  • Experiment with RL, RLHF, RLAIF, reward shaping, policy optimization, and agent training techniques.
  • Document what works, what fails, and why, so we can compound our learnings over time.
  • Stay close to the frontier of LLMs, agents, evaluations, and applied AI engineering.

Connaissances

LLM fine-tuning
Python
PyTorch
Reinforcement learning
RL techniques
Reward modeling
Agent environments

Description du poste

About the AI Studio

The AI Studio's mission is to find the fastest possible path to an autonomous supply chain. We're developing AI agents, learning systems, training models, and more to overcome the biggest challenges remaining in the global supply chain. In short, we are having a lot of fun.

Your Mission In This Role

We're looking for an ambitious ML Engineer focused on LLMs, agents, and reinforcement learning to help build the training, evaluation, and tooling systems behind robust AI decision-making products. You'll work across LLM fine-tuning, agent environments, reward modeling, evaluations, data pipelines, and AI workflow tooling. The role is hands-on: designing experiments, shipping production code, improving model behaviour, and building the infrastructure that lets us learn quickly from both automated and human feedback. You'll help shape how we use LLMs inside agentic systems, how we evaluate model and agent performance, and how we turn feedback into better training data and better behaviour. This role requires mandatory RL training experience with LLMs, including designing and iterating on rewards, reviewing LLM traces, identifying reward hacking or shortcut behaviour, and understanding when the reward signal, environment, or training process needs to change.

Responsibilities
  • Design and implement LLM-powered agent environments for supply chain decision-making
  • Fine-tune, adapt, and evaluate LLMs for domain-specific reasoning and decision support
  • Design, test, and iterate on reward functions that capture the behaviors we want from LLM agents
  • Review LLM traces and rollouts to understand model reasoning, failure modes, reward hacking, and shortcut behaviour
  • Identify when an LLM is exploiting the reward function, escaping the intended RL process, or optimizing for proxy metrics instead of the real objective
  • Improve reward models, environment design, prompts, tools, and feedback loops based on observed model behaviour
  • Build evaluation frameworks to measure model quality, agent performance, robustness, and failure modes
  • Create data pipelines for training, fine-tuning, preference data, synthetic data generation, and human feedback collection
  • Develop tooling that improves how the team builds, tests, debugs, and deploys AI-assisted workflows
  • Experiment with RL, RLHF, RLAIF, reward shaping, policy optimization, and agent training techniques
  • Document what works, what fails, and why, so we can compound our learnings over time
  • Stay close to the frontier of LLMs, agents, evaluations, and applied AI engineering
We want to talk if you:
  • You've trained or fine-tuned LLMs
  • Are excited about AI-assisted tools and getting the most out of them
  • Build & customize your own AI workflows
  • Have experience working with AI agents and RL environments in production
  • Are proficient in Python and PyTorch
  • Can balance research exploration with shipping working code
  • Hands on experience with RL techniques (reward shaping, policy optimization, RLHF)
  • Thrive in fast-moving environments where priorities shift
  • Care about craft in your work
  • Are curious about why things work, not just that they work
Bonus points if:
  • You have experience with human-in-the-loop ML systems
  • You’ve built evaluation frameworks for open-ended tasks
  • You’re familiar with supply chain, logistics, or operations domains
  • You have a side project that shows you can't stop tinkering
Our Values

Core Values All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status.

A Culture, Not just a Company Our culture is built on collaboration, respect and having fun whenever possible. This along with ensuring we live our Core Values, enables us to learn, grow, create customer value, and manage the well‑being of our collective ecosystem of partners, customers and associates alike. We rise each day with an unwavering commitment to lifelong learning, collaborating with respect, and operating with integrity.

What do we do?

Blue Yonder is the AI company for supply chain. As the world leader in end-to-end digital supply chain transformation, Blue Yonder offers a unified, AI-driven platform and multi-tier network that empowers businesses to operate sustainably, scale profitably, and delight their customers—at machine speed. A pioneer in applying AI solutions to the most complicated supply chain challenges, Blue Yonder’s modern innovations and unmatched industry expertise help more than 3,000 retailers, manufacturers, and logistics service providers confidently navigate supply chain complexity and disruption.

"Blue Yonder" is a trademark or registered trademark of Blue Yonder Group, Inc. Any trade, product or service name referenced in this document using the name "Blue Yonder" is a trademark and/or property of Blue Yonder Group, Inc. Blue Yonder, Inc. - Administrative Office 15059 N Scottsdale Rd Suite 350 Scottsdale, AZ 85254-2666

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Staff AI Engineer
Staff AI Engineer

JDA Software • Paris

Sur place
EUR 120 000 - 180 000
AI Engineer
AI Engineer

JDA Software • Paris

Sur place
EUR 70 000 - 110 000
Director, Reinforcement Learning & Agentic Post-Training
Director, Reinforcement Learning & Agentic Post-Training

Blue Yonder • Paris

Sur place
EUR 90 000 - 120 000
Machine Learning Engineer - Reinforcement Learning
Machine Learning Engineer - Reinforcement Learning

Blue Yonder • Paris

Sur place
EUR 90 000 - 120 000
Staff AI Engineer
Staff AI Engineer

Blue Yonder • Paris

Sur place
EUR 90 000 - 140 000
Director, Model Behavior & Evaluation Systems
Director, Model Behavior & Evaluation Systems

Blue Yonder • Paris

Sur place
EUR 140 000 - 200 000
RL & LLMs Engineer for Autonomous Supply Chain
RL & LLMs Engineer for Autonomous Supply Chain

JDA Software • Paris

Sur place
EUR 70 000 - 110 000
AI Engineer
AI Engineer

Blue Yonder • Paris

Sur place
EUR 85 000 - 110 000
Director, Agentic RL for Autonomous Supply Chains
Director, Agentic RL for Autonomous Supply Chains

Blue Yonder • Paris

Sur place
EUR 90 000 - 120 000
Staff AI Engineer: Build Scalable AI Agent Platform
Staff AI Engineer: Build Scalable AI Agent Platform

Blue Yonder • Paris

Sur place
EUR 90 000 - 140 000