Director of AI Agent Behavior & Evaluation

JDA Software

Paris

Sur place

EUR 150 000 - 190 000

Plein temps

14 jours+
Générateur de candidature

Démarquez-vous pour ce poste — générez un CV et une lettre de motivation personnalisés en environ une minute.

Passez les filtres ATS

Résumé du poste

Blue Yonder seeks a deeply technical Director of Model Behavior & Evaluation Systems to own the behavioural quality system for its LLM agents. You will define what good means, establish release gates, and build feedback loops turning traces, SME review, telemetry, and eval failures into model improvements.

This director-level role requires leadership through systems, standards, and people, guiding senior engineers toward robust, scalable infrastructure while aligning with customer workflows in

Qualifications

  • Experience leading technical work on LLM products, AI agents, model behaviour, model evaluation, post-training, AI alignment, or AI product quality.

Responsabilités

  • Own the behavioural quality bar for Blue Yonder's LLM agents across customer-facing supply chain workflows.
  • Build and lead the model behaviour and evaluation systems function across behaviour specs, eval governance, SME review, release gates, regression coverage, and launch readiness.
  • Define launch criteria across operational correctness, tool-use accuracy, workflow completion, escalation quality, safe fallback behaviour, refusal quality, consistency, and customer trust.
  • Establish evaluation authority so evals become release decision infrastructure, not just model-quality reporting.
  • Set the technical direction for behaviour and eval infrastructure across Python eval harnesses, OpenAI Agents SDK workflows, Langfuse traces, LLM-as-judge workflows, deterministic checks, trace analysis, reward/report versioning, and model-candidate comparison.
  • Convert model traces, tool-call failures, SME feedback, red-team findings, telemetry, and customer-facing failures into behaviour specs, eval requirements, training data needs, and model improvement priorities.
  • Partner with the reinforcement learning and post-training organization to turn behaviour gaps into SFT data, preference data, reward criteria, curriculum, NeMo RL experiments, model-candidate decisions, and regression tests.
  • Partner with workflow, data, product, and domain experts to turn supply chain workflow truth into durable scenario coverage, rubrics, synthetic scenarios, eval datasets, and training data requirements.
  • Partner with agent architecture and product engineering teams to ensure prompts, tools, APIs, skills, system instructions, and product workflows express the intended model behaviour consistently.
  • Review model traces, eval outputs, experiment summaries, dataset slices, reward reports, and post-training results closely enough to make informed launch and roadmap decisions.
  • Own customer and user behaviour discovery for agent workflows: what users expect agents to do, explain, ask, verify, elevate, refuse, and act on.
  • Lead red-teaming and behavioural risk programmes for hallucinated operational claims, incorrect tool use, overconfidence, poor escalation, prompt injection, unsafe recommendations, sycophancy, and unhelpful refusals.
  • Define the operating cadence for behavioural quality reviews, model release readiness, model cards, issue triage, regression management, and post-launch behaviour monitoring.
  • Build a high-performing team or function around model behaviour, evaluation governance, SME review systems, and behavioural launch quality.
  • Communicate model behaviour strategy, quality tradeoffs, launch readiness, and residual risks clearly to executives, customers, product teams, engineering teams, and domain experts.

Connaissances

LLM products
AI agents
model behaviour
model evaluation
post-training
AI alignment
AI product quality
leadership
cross-functional collaboration

Outils

Python
PyTorch
Hugging Face Transformers
Hugging Face Datasets
OpenAI Agents SDK
Langfuse
LLM evaluation harnesses
tool-calling traces
model checkpoints
production observability

Description du poste

Blue Yonder seeks a deeply technical Director of Model Behavior & Evaluation Systems to own the behavioural quality system for its LLM agents. You will define what good means, establish release gates, and build feedback loops turning traces, SME review, telemetry, and eval failures into model improvements.

This director-level role requires leadership through systems, standards, and people, guiding senior engineers toward robust, scalable infrastructure while aligning with customer workflows in

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Head of LLM Behavior & Evaluation Systems
Head of LLM Behavior & Evaluation Systems

Blue Yonder • Paris

Sur place
EUR 150 000 - 190 000
Director, Model Behavior & Evaluation Systems
Director, Model Behavior & Evaluation Systems

Blue Yonder • Paris

Sur place
EUR 150 000 - 190 000
Director, Model Behavior & Evaluation Systems
Director, Model Behavior & Evaluation Systems

JDA Software • Paris

Sur place
EUR 150 000 - 190 000
Director, Reinforcement Learning & Agentic Post-Training
Director, Reinforcement Learning & Agentic Post-Training

Blue Yonder • Paris

Sur place
EUR 90 000 - 120 000
Model Behavior Architect- Safety
Model Behavior Architect- Safety

Mistral • Paris

Sur place
EUR 70 000 - 90 000
Director, Agentic RL for Autonomous Supply Chains
Director, Agentic RL for Autonomous Supply Chains

Blue Yonder • Paris

Sur place
EUR 90 000 - 120 000
LLM Behavior Architect: Shape Next‑Gen AI
LLM Behavior Architect: Shape Next‑Gen AI

Mistral • Paris

Sur place
EUR 70 000 - 90 000
Remote LLM Evaluator & Model Response Analyst
Remote LLM Evaluator & Model Response Analyst

Odixcity Consulting • La Réunion

Sur place
EUR 52 000 - 78 000
Founding Staff AI Engineer — LLM Fine-Tuning & Data
Founding Staff AI Engineer — LLM Fine-Tuning & Data

BAO • Paris

Hybride
EUR 120 000 - 180 000
AI Software Engineer: Lead LLM Strategy & Agents
AI Software Engineer: Lead LLM Strategy & Agents

Jobtailor • Grenoble

Sur place
EUR 60 000 - 90 000