Head of LLM Behavior & Evaluation Systems

Blue Yonder

Paris

Sur place

EUR 150 000 - 190 000

Plein temps

14 jours+
Générateur de candidature

Obtenez une réponse de cet employeur — un CV et une lettre de motivation adaptés exactement à ce qu’il recherche.

Passez les filtres ATS

Résumé du poste

Blue Yonder is seeking a deeply technical Director of Model Behavior & Evaluation Systems to own the behavioral quality system for its LLM agents. This leader will define what good behavior means, establish release gates, and build feedback loops from traces, SME review, telemetry, and eval failures into model improvements.

The role demands systems thinking, technical depth in the modern ML stack, and the ability to guide senior engineers and cross-functional partners toward robust, scalable

Qualifications

  • Led technical work on LLM products, AI agents, or AI product quality.
  • Built evaluation systems for open-ended LLM behavior, including rubrics and release gates.
  • Hands-on depth with LLM tool calling, function calling, or API agents.
  • Fluent in modern LLM stack: Python, PyTorch, Hugging Face ecosystems.
  • Understand post-training workflows: supervised fine-tuning, reward modeling, RLHF/RLAIF.
  • Capable of inspecting model traces, eval failures, telemetry, and data to guide fixes.
  • Strong product judgment translating user needs into shipped behavior.

Responsabilités

  • Own the behavioral quality bar for Blue Yonder's LLM agents across workflows.
  • Define launch criteria across operational correctness, tool-use, and escalation quality.
  • Establish evaluation authority and release gates for deployments.
  • Set technical direction for eval infra: Python harnesses, Langfuse traces, API workflows.
  • Convert traces, SME feedback, telemetry into specs, data needs, and improvements.
  • Partner with RL and post-training org to shape data and experiments.
  • Collaborate with domain experts to cover supply chain scenarios.
  • Lead red-teaming and behavioral risk programs for unsafe behavior and misuses.
  • Define cadence for behavioral reviews, model cards, and post-launch monitoring.
  • Build a high-performing team around model behavior and evaluation governance.
  • Communicate strategy, risks, and decisions to executives and teams.

Connaissances

LLM product leadership
Evaluation systems
Tool calling
Python stack
Post-training workflows
Failure analysis
Trace inspection
Product judgment
Cross-functional leadership
Launch readiness
Excellent communicator
Model specs
Senior teams

Outils

NVIDIA NeMo RL
OpenAI Agents SDK
Langfuse
Hugging Face Transformers
Hugging Face Datasets
PyTorch
Python

Description du poste

Blue Yonder is seeking a deeply technical Director of Model Behavior & Evaluation Systems to own the behavioral quality system for its LLM agents. This leader will define what good behavior means, establish release gates, and build feedback loops from traces, SME review, telemetry, and eval failures into model improvements.

The role demands systems thinking, technical depth in the modern ML stack, and the ability to guide senior engineers and cross-functional partners toward robust, scalable

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Director of AI Agent Behavior & Evaluation
Director of AI Agent Behavior & Evaluation

JDA Software • Paris

Sur place
EUR 150 000 - 190 000
Director, Model Behavior & Evaluation Systems
Director, Model Behavior & Evaluation Systems

Blue Yonder • Paris

Sur place
EUR 150 000 - 190 000
Director, Model Behavior & Evaluation Systems
Director, Model Behavior & Evaluation Systems

JDA Software • Paris

Sur place
EUR 150 000 - 190 000
LLM Behavior Architect: Shape Next‑Gen AI
LLM Behavior Architect: Shape Next‑Gen AI

Mistral • Paris

Sur place
EUR 70 000 - 90 000
Model Behavior Architect- Safety
Model Behavior Architect- Safety

Mistral • Paris

Sur place
EUR 70 000 - 90 000
Remote LLM Evaluator & Model Response Analyst
Remote LLM Evaluator & Model Response Analyst

Odixcity Consulting • La Réunion

Sur place
EUR 52 000 - 78 000
LLM Evaluator (Model Response Analyst)
LLM Evaluator (Model Response Analyst)

Odixcity Consulting • La Réunion

Sur place
EUR 52 000 - 78 000
Senior LLM Evaluation Engineer & Coding Architect (Remote)
Senior LLM Evaluation Engineer & Coding Architect (Remote)

Braintrust • Paris

Hybride
EUR 110 000 - 150 000
Director, Reinforcement Learning & Agentic Post-Training
Director, Reinforcement Learning & Agentic Post-Training

Blue Yonder • Paris

Sur place
EUR 90 000 - 120 000
AI Engineer: LLM & Agentic Systems Orchestrator
AI Engineer: LLM & Agentic Systems Orchestrator

Jobtailor • Paris

Sur place
EUR 60 000 - 90 000