Director, Model Behavior & Evaluation Systems

jda

Paris

Sur place

EUR 120 000 - 180 000

Plein temps

Il y a 5 jours
Soyez parmi les premiers à postuler
Générateur de candidature

Obtenez une réponse de cet employeur — un CV et une lettre de motivation adaptés exactement à ce qu’il recherche.

Passez les filtres ATS

Résumé du poste

Blue Yonder seeks a deeply technical Director of Model Behavior & Evaluation Systems to own the behavioral quality system for our LLM agents operating in supply chain workflows. You will define what good means for agents, establish release gates, and build feedback loops from traces, SME review, telemetry, and eval failures into improvements.

You will lead through systems, standards, people, and decisions, collaborating with senior engineers and product partners to shape the end-to-end

Qualifications

  • Director-level technical leadership role in AI for supply chain.

Responsabilités

  • Own the behavioral quality bar for Blue Yonder's LLM agents across customer-facing workflows.
  • Build and lead the model behavior and evaluation systems function across specs, governance, release gates, regression coverage, and launch readiness.
  • Define launch criteria across operational correctness, tool-use accuracy, workflow completion, escalation quality, safe fallback behavior, refusal quality, consistency, and customer trust.
  • Establish evaluation authority so evals become release decision infrastructure, not just reporting.
  • Set technical direction for behavior and eval infra across Python eval harnesses, SDK workflows, Langfuse traces, and deterministic checks.
  • Convert traces, tool-call failures, SME feedback, red-team findings, telemetry, and customer-facing failures into behavior specs and training data needs.
  • Partner with RL and post-training orgs to turn gaps into SFT data, preference data, reward criteria, curriculum, NeMo RL experiments, and regression tests.
  • Partner with workflow, data teams to align on data-ready processes and tooling.

Connaissances

Leadership
System thinking
ML model evaluation
Stakeholder collaboration

Outils

PyTorch
Hugging Face Transformers
Langfuse
OpenAI Agents SDK
NVIDIA NeMo RL
LLM evaluation harnesses

Description du poste

About Blue Yonder

Blue Yonder is the AI company for supply chain. Our platform helps the world's leading companies plan, fulfill, deliver, and operate more resilient supply chains across complex global networks.

We are building toward the autonomous supply chain: intelligent systems that can understand operational context, reason through tradeoffs, use tools, collaborate with people, and take action across real supply chain workflows.

About Autonomy Labs

Autonomy Labs' mission is to find the fastest possible path to an autonomous supply chain.

We build LLM agents, learning systems, model training pipelines, evaluations, simulations, and decision-making systems for some of the hardest problems in global supply chain. The work spans LLMs, agentic workflows, tool use, software automation, evaluation, post-training, optimization, and production engineering.

In short, we are having a lot of fun.

Your Mission

We are looking for a deeply technical Director of Model Behavior & Evaluation Systems to own the behavioral quality system for Blue Yonder's LLM agents.

Our agents are not generic chatbots. They are being trained to operate supply chain software: querying state, calling APIs, interpreting operational context, proposing actions, handling exceptions, asking for missing information, and helping users make decisions in complex enterprise environments.

Your mission is to make model behavior a product-quality system, not a collection of dashboards. You will define what \"good\" means for agents operating supply chain workflows, establish the release gates that determine when behavior is ready to ship, and build the feedback loops that turn traces, customer feedback, SME review, telemetry, red-teaming, and eval failures into model improvements.

This is a director-level technical leadership role. You will lead through systems, standards, people, and decisions. You should be close enough to model traces, evals, post-training, tool use, and customer workflows to make strong technical calls, while operating at the level of ownership boundaries, launch authority, roadmap sequencing, and team building.

The stack is real and close to the work. You should expect to operate around Python, PyTorch, Hugging Face Transformers and Datasets, NVIDIA NeMo RL, OpenAI Agents SDK, Langfuse, LLM evaluation harnesses, tool-calling traces, model checkpoints, reward and preference data, synthetic scenarios, experiment reports, and production observability. You do not need to be the person implementing every pipeline, but you do need the technical depth to challenge designs, read artifacts, understand failure modes, and guide senior engineers toward better systems.

What You'll Do:
  • Own the behavioral quality bar for Blue Yonder's LLM agents across customer-facing supply chain workflows.

  • Build and lead the model behavior and evaluation systems function across behavior specs, eval governance, SME review, release gates, regression coverage, and launch readiness.

  • Define launch criteria across operational correctness, tool-use accuracy, workflow completion, escalation quality, safe fallback behavior, refusal quality, consistency, and customer trust.

  • Establish evaluation authority so evals become release decision infrastructure, not just model-quality reporting.

  • Set the technical direction for behavior and eval infrastructure across Python eval harnesses, OpenAI Agents SDK workflows, Langfuse traces, LLM-as-judge workflows, deterministic checks, trace analysis, reward/report versioning, and model-candidate comparison.

  • Convert model traces, tool-call failures, SME feedback, red-team findings, telemetry, and customer-facing failures into behavior specs, eval requirements, training data needs, and model improvement priorities.

  • Partner with the reinforcement learning and post-training organization to turn behavior gaps into SFT data, preference data, reward criteria, curriculum, NeMo RL experiments, model-candidate decisions, and regression tests.

  • Partner with workflow, data

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Director, Model Behavior & Evaluation Systems
Director, Model Behavior & Evaluation Systems

Blue Yonder • Paris

Sur place
EUR 140 000 - 200 000
Director, AI Agent Behavior & Evaluation Systems
Director, AI Agent Behavior & Evaluation Systems

Blue Yonder • Paris

Sur place
EUR 140 000 - 200 000
Head of AI Behavior & Evaluation Systems
Head of AI Behavior & Evaluation Systems

jda • Paris

Sur place
EUR 120 000 - 180 000
Director, Reinforcement Learning & Agentic Post-Training
Director, Reinforcement Learning & Agentic Post-Training

Blue Yonder • Paris

Sur place
EUR 90 000 - 120 000
Director, Reinforcement Learning & Agentic Post-Training
Director, Reinforcement Learning & Agentic Post-Training

jda • Paris

Sur place
EUR 120 000 - 180 000
Staff AI Engineer
Staff AI Engineer

Blue Yonder • Paris

Sur place
EUR 90 000 - 140 000
Staff AI Engineer
Staff AI Engineer

jda • Paris

Sur place
EUR 120 000 - 180 000
Staff AI Engineer
Staff AI Engineer

JDA Software • Paris

Sur place
EUR 120 000 - 180 000
AI Engineer
AI Engineer

jda • Paris

Sur place
EUR 70 000 - 100 000
Machine Learning Engineer - Reinforcement Learning
Machine Learning Engineer - Reinforcement Learning

JDA Software • Paris

Sur place
EUR 70 000 - 110 000