Director, Model Behavior & Evaluation Systems

JDA Software

Paris

Sur place

EUR 150 000 - 190 000

Plein temps

14 jours+
Générateur de candidature

Démarquez-vous pour ce poste — générez un CV et une lettre de motivation personnalisés en environ une minute.

Passez les filtres ATS

Résumé du poste

Blue Yonder seeks a deeply technical Director of Model Behavior & Evaluation Systems to own the behavioural quality system for its LLM agents. You will define what good means, establish release gates, and build feedback loops turning traces, SME review, telemetry, and eval failures into model improvements.

This director-level role requires leadership through systems, standards, and people, guiding senior engineers toward robust, scalable infrastructure while aligning with customer workflows in

Qualifications

  • Experience leading technical work on LLM products, AI agents, model behaviour, model evaluation, post-training, AI alignment, or AI product quality.

Responsabilités

  • Own the behavioural quality bar for Blue Yonder's LLM agents across customer-facing supply chain workflows.
  • Build and lead the model behaviour and evaluation systems function across behaviour specs, eval governance, SME review, release gates, regression coverage, and launch readiness.
  • Define launch criteria across operational correctness, tool-use accuracy, workflow completion, escalation quality, safe fallback behaviour, refusal quality, consistency, and customer trust.
  • Establish evaluation authority so evals become release decision infrastructure, not just model-quality reporting.
  • Set the technical direction for behaviour and eval infrastructure across Python eval harnesses, OpenAI Agents SDK workflows, Langfuse traces, LLM-as-judge workflows, deterministic checks, trace analysis, reward/report versioning, and model-candidate comparison.
  • Convert model traces, tool-call failures, SME feedback, red-team findings, telemetry, and customer-facing failures into behaviour specs, eval requirements, training data needs, and model improvement priorities.
  • Partner with the reinforcement learning and post-training organization to turn behaviour gaps into SFT data, preference data, reward criteria, curriculum, NeMo RL experiments, model-candidate decisions, and regression tests.
  • Partner with workflow, data, product, and domain experts to turn supply chain workflow truth into durable scenario coverage, rubrics, synthetic scenarios, eval datasets, and training data requirements.
  • Partner with agent architecture and product engineering teams to ensure prompts, tools, APIs, skills, system instructions, and product workflows express the intended model behaviour consistently.
  • Review model traces, eval outputs, experiment summaries, dataset slices, reward reports, and post-training results closely enough to make informed launch and roadmap decisions.
  • Own customer and user behaviour discovery for agent workflows: what users expect agents to do, explain, ask, verify, elevate, refuse, and act on.
  • Lead red-teaming and behavioural risk programmes for hallucinated operational claims, incorrect tool use, overconfidence, poor escalation, prompt injection, unsafe recommendations, sycophancy, and unhelpful refusals.
  • Define the operating cadence for behavioural quality reviews, model release readiness, model cards, issue triage, regression management, and post-launch behaviour monitoring.
  • Build a high-performing team or function around model behaviour, evaluation governance, SME review systems, and behavioural launch quality.
  • Communicate model behaviour strategy, quality tradeoffs, launch readiness, and residual risks clearly to executives, customers, product teams, engineering teams, and domain experts.

Connaissances

LLM products
AI agents
model behaviour
model evaluation
post-training
AI alignment
AI product quality
leadership
cross-functional collaboration

Outils

Python
PyTorch
Hugging Face Transformers
Hugging Face Datasets
OpenAI Agents SDK
Langfuse
LLM evaluation harnesses
tool-calling traces
model checkpoints
production observability

Description du poste

About Blue Yonder

Blue Yonder is the AI company for supply chain. Our platform helps the world\'s leading companies plan, fulfill, deliver, and operate more resilient supply chains across complex global networks. We are building toward the autonomous supply chain: intelligent systems that can understand operational context, reason through tradeoffs, use tools, collaborate with people, and take action across real supply chain workflows.



About Autonomy Labs

Autonomy Labs\' mission is to find the fastest possible path to an autonomous supply chain. We build LLM agents, learning systems, model training pipelines, evaluations, simulations, and decision-making systems for some of the hardest problems in global supply chain. The work spans LLMs, agentic workflows, tool use, software automation, evaluation, post-training, optimization, and production engineering. In short, we are having a lot of fun.



Your Mission

We are looking for a deeply technical Director of Model Behavior & Evaluation Systems to own the behavioural quality system for Blue Yonder\'s LLM agents. Our agents are not generic chatbots. They are being trained to operate supply chain software: querying state, calling APIs, interpreting operational context, proposing actions, handling exceptions, asking for missing information, and helping users make decisions in complex enterprise environments.


Your mission is to make model behavior a product-quality system, not a collection of dashboards. You will define what \"good\" means for agents operating supply chain workflows, establish the release gates that determine when behavior is ready to ship, and build the feedback loops that turn traces, customer feedback, SME review, telemetry, red-teaming, and eval failures into model improvements.


This is a director-level technical leadership role. You will lead through systems, standards, people, and decisions. You should be close enough to model traces, evals, post-training, tool use, and customer workflows to make strong technical calls, while operating at the level of ownership boundaries, launch authority, roadmap sequencing, and team building.


The stack is real and close to the work. You should expect to operate around Python, PyTorch, Hugging Face Transformers and Datasets, NVIDIA NeMo RL, OpenAI Agents SDK, Langfuse, LLM evaluation harnesses, tool-calling traces, model checkpoints, reward and preference data, synthetic scenarios, experiment reports, and production observability.


You do not need to be the person implementing every pipeline, but you do need the technical depth to challenge designs, read artifacts, understand failure modes, and guide senior engineers toward better systems.



What You\'ll Do


  • Own the behavioural quality bar for Blue Yonder\'s LLM agents across customer-facing supply chain workflows.

  • Build and lead the model behaviour and evaluation systems function across behaviour specs, eval governance, SME review, release gates, regression coverage, and launch readiness.

  • Define launch criteria across operational correctness, tool-use accuracy, workflow completion, escalation quality, safe fallback behaviour, refusal quality, consistency, and customer trust.

  • Establish evaluation authority so evals become release decision infrastructure, not just model-quality reporting.

  • Set the technical direction for behaviour and eval infrastructure across Python eval harnesses, OpenAI Agents SDK workflows, Langfuse traces, LLM-as-judge workflows, deterministic checks, trace analysis, reward/report versioning, and model-candidate comparison.

  • Convert model traces, tool-call failures, SME feedback, red-team findings, telemetry, and customer-facing failures into behaviour specs, eval requirements, training data needs, and model improvement priorities.

  • Partner with the reinforcement learning and post-training organization to turn behaviour gaps into SFT data, preference data, reward criteria, curriculum, NeMo RL experiments, model-candidate decisions, and regression tests.

  • Partner with workflow, data, product, and domain experts to turn supply chain workflow truth into durable scenario coverage, rubrics, synthetic scenarios, eval datasets, and training data requirements.

  • Partner with agent architecture and product engineering teams to ensure prompts, tools, APIs, skills, system instructions, and product workflows express the intended model behaviour consistently.

  • Review model traces, eval outputs, experiment summaries, dataset slices, reward reports, and post-training results closely enough to make informed launch and roadmap decisions.

  • Own customer and user behaviour discovery for agent workflows: what users expect agents to do, explain, ask, verify, elevate, refuse, and act on.

  • Lead red-teaming and behavioural risk programmes for hallucinated operational claims, incorrect tool use, overconfidence, poor escalation, prompt injection, unsafe recommendations, sycophancy, and unhelpful refusals.

  • Define the operating cadence for behavioural quality reviews, model release readiness, model cards, issue triage, regression management, and post-launch behaviour monitoring.

  • Build a high-performing team or function around model behaviour, evaluation governance, SME review systems, and behavioural launch quality.

  • Communicate model behaviour strategy, quality tradeoffs, launch readiness, and residual risks clearly to executives, customers, product teams, engineering teams, and domain experts.



What We\'re Looking For

We want to talk if you:



  • Have led technical work on LLM products, AI agents, model behaviour, model evaluation, post-training, AI alignment, or AI product quality.

  • Have built or led evaluation systems for open-ended LLM behaviour, including eval datasets, graders, rubrics, scorecards, SME review loops, regression tests, or launch gates.

  • Have hands-on depth with LLM tool calling, function calling, API agents, workflow agents, software-operating agents, or enterprise automation.

  • Are technically fluent in the modern LLM stack, including Python, PyTorch, Hugging Face Transformers, Hugging Face Datasets, OpenAI Agents SDK, Langfuse, model checkpoints, tokenization, inference behaviour, and experiment analysis.

  • Understand post-training workflows well enough to partner deeply with ML teams, including supervised fine-tuning, preference data, RLHF/RLAIF, reward modelling, NeMo RL, system prompting, synthetic data, and model launch evaluation.

  • Can reason about practical LLM behaviour failures: hallucination, sycophancy, overconfidence, verbosity, refusal behaviour, instruction hierarchy, prompt injection, tool-use failures, eval drift, and behavioural regressions.

  • Have enough technical fluency to inspect model traces, tool-call transcripts, eval failures, telemetry, datasets, experiment configs, and model-quality reports, then guide teams toward concrete fixes.

  • Have strong product judgement and experience translating user or customer needs into shipped product behaviour.

  • Can lead cross‑functional work across product, research, engineering, design, security, legal, customer success, domain experts, and executives.

  • Are comfortable making launch-readiness decisions under ambiguity, with explicit trade‑offs and evidence.

  • Are an exceptional written communicator who can turn ambiguous behavioural principles into clear guidelines, examples, rubrics, specs, and decisions.

  • Have experience defining model specifications, behaviour guidelines, system instruction hierarchies, annotation rubrics, policy taxonomies, or model behaviour standards.

  • Have managed, built, or strongly influenced senior technical teams.



Experience That Stands Out


  • Strong candidates may also bring experience in one or more of the following areas: owning model behaviour, assistant quality, agent behaviour, LLM product quality, or AI alignment work at a frontier lab, AI product company, or enterprise AI platform.

  • Scaling an evaluation, model-quality, or launch-readiness function from early experimentation into repeatable operating practice.

  • Working with NVIDIA NeMo Tron, vLLM, Ray, distributed evaluation, large-scale inference, or model-serving observability.

  • Applying AI agents to supply chain planning, warehouse management, transportation, logistics, operations research, retail, manufacturing, or enterprise workflow software.

  • Working with strategic enterprise customers on ambiguous, high-impact AI product behaviour, safety, or launch-readiness decisions.

  • Publishing, open-sourcing, or internally leading notable work around model behaviour, evals, agents, post-training, or AI product quality.



Technical Environment

You should expect to work around: Python, PyTorch, Hugging Face Transformers, Hugging Face Datasets, and modern LLM post-training workflows. NVIDIA NeMo RL, preference data, reward criteria, synthetic scenarios, model candidates, and experiment reports. Tool-calling agents, OpenAI Agents SDK, API traces, workflow state, prompt and system-instruction behaviour, and software-operating agent failure modes. LLM evaluation harnesses, deterministic checks, LLM judges, SME trace review, grader rubrics, regression suites, and release gates. Langfuse and production observability for model behaviour: traces, telemetry, customer feedback, incident review, post-launch regressions, and model-quality reporting.



What Makes This Role Different?

Model behaviour is product behaviour. For supply chain agents, the difference between a useful system and an unsafe or frustrating one often comes down to judgement: when to ask for clarification, when to call a tool, when to explain a tradeoff, when to surface uncertainty, when to elevate to a human, and when not to act. This role owns that judgement as an operating system: the standards, evidence, release gates, teams, and feedback loops that determine whether our agents are ready to operate in real customer workflows.



Equal Opportunity

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or any other legally protected characteristic.



Our Values

If you want to know the heart of a company, take a look at their values. Ours unite us. They are what drive our success – and the success of our customers. Does your heart beat like ours? Find out here: Core Values All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status.



Who are we?

We are a proven, passionate bunch of disruptors. Our work is all about tapping into your potential so we can deliver the best solutions and customer experiences on the planet. Collaboration, respect, and a great work-life balance earned us the title of \"Best Place to Work- Employees\' Choice\" by Glassdoor. Our people are smart, creative, rock stars with over 400 patents and 10,000 people years of domain expertise.



What do we do?

Blue Yonder is the world leader in digital supply chain and omni-channel commerce fulfillment. Our intelligent, end-to-end platform enables retailers, manufacturers and logistics providers to seamlessly predict, pivot and fulfil customer demand. With Blue Yonder, you can make more automated, profitable business decisions that deliver greater growth and re-imagined customer experiences.



Blue Yonder - Fulfill your Potential.

™ blueyonder.com



Trademark Information

“Blue Yonder” is a trademark or registered trademark of Blue Yonder, Inc. Any trade, product or service name referenced in this document using the name “Blue Yonder” is a trademark and/or property of Blue Yonder, Inc.


Blue Yonder, Inc. 15059 N Scottsdale Rd, Ste 400 Scottsdale, AZ 85254

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Director, Model Behavior & Evaluation Systems
Director, Model Behavior & Evaluation Systems

Blue Yonder • Paris

Sur place
EUR 150 000 - 190 000
Director, Reinforcement Learning & Agentic Post-Training
Director, Reinforcement Learning & Agentic Post-Training

Blue Yonder • Paris

Sur place
EUR 90 000 - 120 000
Director of AI Agent Behavior & Evaluation
Director of AI Agent Behavior & Evaluation

JDA Software • Paris

Sur place
EUR 150 000 - 190 000
Head of LLM Behavior & Evaluation Systems
Head of LLM Behavior & Evaluation Systems

Blue Yonder • Paris

Sur place
EUR 150 000 - 190 000
Head of AI Enablement
Head of AI Enablement

Bigblue • Paris

Sur place
EUR 80 000 - 120 000
10€ meal voucher
Unlimited snacks
100% health insurance coverage
+1
Head of AI Enablement
Head of AI Enablement

Bigblue • Paris

Sur place
EUR 80 000 - 100 000
Meal vouchers
Unlimited snacks
ClassPass membership
+1
Senior Data Labeler
Senior Data Labeler

Whitecircle • Paris

Hybride
EUR 90 000 - 150 000
Competitive salary + equity
Hybrid work Paris/London
Relocation package (Paris)
+4
Senior Data Labeler
Senior Data Labeler

Visa Hunt • Paris

Hybride
EUR 65 000 - 90 000
Competitive salary + equity
Hybrid work from Paris with relocation
London option (limited relocation/ins)
+3
Senior Applied Scientist - AI Platform
Senior Applied Scientist - AI Platform

Datadog • Paris

Sur place
EUR 120 000 - 160 000
RSUs
ESPP
Professional development
+3
ML Infrastructure Engineer
ML Infrastructure Engineer

Whitecircle • Paris

Sur place
EUR 90 000 - 140 000