LLM Evaluator (Model Response Analyst)

Odixcity Consulting

La Réunion

Sur place

EUR 52 000 - 78 000

Plein temps

Il y a 5 jours
Soyez parmi les premiers à postuler
Générateur de candidature

Transformez ce poste en entretien — un CV et une lettre de motivation conçus selon ce que cet employeur recherche.

Passez les filtres ATS

Résumé du poste

Odixcity Consulting seeks an experienced LLM Evaluator to assess, analyze, and improve large language model performance. You will evaluate AI-generated content for factuality, coherence, safety, and alignment with guidelines.

The role involves ranking outputs, providing justified choices, and reporting recurring failures to help patch vulnerabilities. You will collaborate with QA to refine evaluation guidelines, participate in cross-checking sessions, and explore deeper causes behind errors to

Qualifications

  • Minimum of 2 years in a relevant field such as NLP/AI or data analysis.
  • Bachelor’s degree in Computer Science or related field.
  • Deep understanding of how prompts influence model behavior.
  • Ability to explain why outputs are good or bad based on criteria like factuality and adherence.
  • Experience with RLHF data collection.
  • Experience monitoring IA A scores and calibrating judgments.
  • Experience sourcing, cleaning, and annotating datasets for eval.
  • Familiarity with A/B testing concepts for model comparison.

Responsabilités

  • Evaluate and rank model-generated text using complex rubrics for factuality, coherence, safety, and instruction adherence.
  • Review multiple model outputs and justify preferred choices.
  • Provide actionable feedback to modeling/training teams on failing models.
  • Collaborate in cross-checking sessions to calibrate scoring across a global team.
  • Investigate underlying causes of errors and hypothesize reasons like data issues or prompt misinterpretation.
  • Flag novel model behaviors to researchers for a living library of failure modes.

Connaissances

NLP data analysis
Prompt engineering
RLHF data collection
IAA analysis
Quality Assurance
Technical writing
Data cleaning
Experiment design
Calibration sessions

Formation

Bachelor's in Computer Science

Outils

Python
Jupyter
NLP toolkit

Description du poste

Job Title: LLM Evaluator (Model Response Analyst)

Location: Remote (Worldwide)

Job Summary: We are seeking a detail-oriented and analytical LLM Evaluator to assess, analyze, and improve the performance of large language models (LLMs). In this role, you will evaluate AI-generated content for accuracy, coherence, factual reliability, bias, safety, and alignment with defined guidelines.

Responsibilities
  • Evaluate and rank model-generated text based on complex rubrics covering dimensions such as factuality, coherence, safety, instruction- following, and creativity.
  • Review multiple model responses to the same prompt and determine which output a human would prefer, providing justifications for your choices.
  • Provide clear, concise feedback to the modeling and training teams regarding recurring failure models observed during evaluation sessions.
  • Attempt to “break” the model by crafting prompts designed to elicit biased, harmful, or insecure outputs to help patch safety vulnerabilities.
  • Collaborate with the quality assurance team to suggest improvements to evaluation guidelines when you encounter ambiguous or unclassifiable edge cases.
  • Participate in regular “cross-checking” sessions with other evaluators to calibrate scoring standards and ensure inter-rater reliability across the global team.
  • When a model underperforms, dig deeper than the surface score to hypothesize “why” the model made a specific error (e.g., training data vs. prompt misinterpretation).
  • Identify and flag novel or unexpected model behaviors to the research team, contributing to a living library of unique model outputs and failure modes.
Requirements
  • Minimum of 2 years of professional experience in a relevant field such as; Computational Linguistics, Data Analysis, Technical Writing, Quality Assurance (specifically for NLP/AI), or cognitive science.
  • Bachelor’s degree in Computer Science, or a relating field.
  • Deep understanding of how-to craft prompts to elicit specific behaviors and test model limits.
  • Ability to look at a text output and explain “why” it is “good” or “bad” based on logic, tone, factuality, and instruction adherence.
  • Experience working with Reinforcement Learning from Human Feedback (RLHF) data collection.
  • Proven experience monitoring and improving consistency among evaluation teams. Ability to analyze IAA scores and conduct calibration sessions to align judgement.
  • Experience sourcing, cleaning, and annotating datasets specifically for the fine-tuning or evaluating LLMs. Understanding of data distribution and its impact on model performance.
  • Familiarity with A/B testing concepts applied to AI. Ability to help design experiments to test if a new model version is truly “better” than the previous one.
Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Remote LLM Evaluator & Model Response Analyst
Remote LLM Evaluator & Model Response Analyst

Odixcity Consulting • La Réunion

Sur place
EUR 52 000 - 78 000
Language Specialist | $50/hr Remote
Language Specialist | $50/hr Remote

Crossing Hurdles • France

Sur place
EUR 32 000 - 52 000
Language Specialist
Language Specialist

Innodata Inc. • France

Sur place
EUR 28 000 - 42 000
Model Behavior Architect- Safety
Model Behavior Architect- Safety

Mistral • Paris

Sur place
EUR 70 000 - 90 000
LLM Engineer
LLM Engineer

Licorne Society • Paris

Sur place
EUR 90 000 - 130 000
AI Evaluation Specialist | $35/hr Remote
AI Evaluation Specialist | $35/hr Remote

Crossing Hurdles • La Réunion

Sur place
EUR 42 000 - 64 000
Remote LLM Evaluation Scientist - Benchmark Frontiers
Remote LLM Evaluation Scientist - Benchmark Frontiers

Anyone AI • Paris

Sur place
EUR 90 000 - 140 000
Lead AI Engineer
Lead AI Engineer

Jobtailor • Paris

Sur place
EUR 90 000 - 140 000
Principal Coding Annotator / LLM Evaluation Engineer
Principal Coding Annotator / LLM Evaluation Engineer

Braintrust • France

Hybride
GBP 70 706 - 141 412
Staff Data Analysis & Evaluation - LLM Robustness (Remote)
Staff Data Analysis & Evaluation - LLM Robustness (Remote)

Cohere • Paris

Sur place
EUR 70 000 - 90 000
Inclusive culture
Weekly lunch stipend
Full health and dental benefits
+5