LLM Evaluator (Model Response Analyst)

Odixcity Consulting

South Africa

Remoto

ZAR 360.000 - 540.000

Tempo pieno

2 giorni fa
Candidati tra i primi
Generatore di candidature

Non inviare un curriculum generico — genera un curriculum e una lettera di presentazione personalizzati per questo specifico impiego.

Supera i filtri ATS

Descrizione del lavoro

Odixcity Consulting seeks a detail-oriented LLM Evaluator to assess and improve large language models. You will analyze model outputs for factuality, coherence, safety, and alignment with guidelines, providing actionable feedback to ML teams.

You will review multiple responses, justify preferences, and help calibrate scoring with global evaluators, contributing to a living library of model behaviors and failure modes.

Competenze

  • Minimum 2 years of professional experience in NLP/AI-related fields.
  • Bachelor’s degree in CS or related field.
  • Ability to explain why text is “good” or “bad” based on logic, tone, factuality and adherence to guidelines.
  • Experience with RLHF data collection.
  • Experience coordinating evaluation teams and IAA calibration.

Mansioni

  • Evaluate and rank model outputs using rubrics covering factuality, coherence, safety, instruction-following, and creativity.
  • Review multiple responses to a prompt and justify preferred output with clear reasoning.
  • Provide feedback to ML teams on recurring failures observed during evaluation.
  • Craft prompts to challenge models and reveal safety vulnerabilities.
  • Collaborate to update evaluation guidelines for edge cases and ambiguities.
  • Participate in cross-checking sessions to calibrate scoring across evaluators.
  • Analyze root causes of errors beyond surface scores.
  • Flag novel model behaviors to researchers for library of outputs and failure modes.

Conoscenze

Computational Linguistics
Data Analysis
Technical Writing
Quality Assurance (NLP/AI)
Cognitive Science
Prompting & Prompt Design
RLHF Data Collection

Formazione

Bachelor’s degree in Computer Science or related field

Descrizione del lavoro

Job Title: LLM Evaluator (Model Response Analyst)

Location: Remote (Worldwide)

Job Summary: We are seeking a detail-oriented and analytical LLM Evaluator to assess, analyze, and improve the performance of large language models (LLMs). In this role, you will evaluate AI-generated content for accuracy, coherence, factual reliability, bias, safety, and alignment with defined guidelines.

Responsibilities
  • Evaluate and rank model-generated text based on complex rubrics covering dimensions such as factuality, coherence, safety, instruction- following, and creativity.
  • Review multiple model responses to the same prompt and determine which output a human would prefer, providing justifications for your choices.
  • Provide clear, concise feedback to the modeling and training teams regarding recurring failure models observed during evaluation sessions.
  • Attempt to "break" the model by crafting prompts designed to elicit biased, harmful, or insecure outputs to help patch safety vulnerabilities.
  • Collaborate with the quality assurance team to suggest improvements to evaluation guidelines when you encounter ambiguous or unclassifiable edge cases.
  • Participate in regular "cross-checking" sessions with other evaluators to calibrate scoring standards and ensure inter-rater reliability across the global team.
  • When a model underperforms, dig deeper than the surface score to hypothesize "why" the model made a specific error (e.g., training data vs. prompt misinterpretation).
  • Identify and flag novel or unexpected model behaviors to the research team, contributing to a living library of unique model outputs and failure modes.
Requirements
  • Minimum of 2 years of professional experience in a relevant field such as; Computational Linguistics, Data Analysis, Technical Writing, Quality Assurance (specifically for NLP/AI), or cognitive science.
  • Bachelor’s degree in Computer Science, or a relating field.
  • Deep understanding of how-to craft prompts to elicit specific behaviors and test model limits.
  • Ability to look at a text output and explain "why" it is "good" or "bad" based on logic, tone, factuality, and instruction adherence.
  • Experience working with Reinforcement Learning from Human Feedback (RLHF) data collection.
  • Proven experience monitoring and improving consistency among evaluation teams. Ability to analyze IAA scores and conduct calibration sessions to align judgement.
  • Experience sourcing, cleaning, and annotating datasets specifically for the fine-tuning or evaluating LLMs. Understanding of data distribution and its impact on model performance.
  • Familiarity with A/B testing concepts applied to AI. Ability to help design experiments to test if a new model version is truly "better" than the previous one.
Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Remote LLM Evaluator: Model Response Analyst
Remote LLM Evaluator: Model Response Analyst

Odixcity Consulting • Sud Africa

Remoto
ZAR 360.000 - 540.000
RLHF Specialist
RLHF Specialist

Odixcity Consulting • Sud Africa

Remoto
ZAR 1.477.000 - 2.626.000
Generative AI Evaluator | $30/hr Remote
Generative AI Evaluator | $30/hr Remote

Crossing Hurdles • Sud Africa

In loco
ZAR 441.892 - 662.838
Remote AI Analyst: Multilingual Research & LLM Tuning
Remote AI Analyst: Multilingual Research & LLM Tuning

Turing • Sud Africa

In loco
ZAR 30.000 - 50.000
Fully remote environment
Work on cutting-edge AI projects
Potential for contract extension based on performance
Generative AI Quality Evaluator - Remote
Generative AI Quality Evaluator - Remote

Crossing Hurdles • Sud Africa

Remoto
ZAR 441.892 - 662.838
Software Engineer Annotator
Software Engineer Annotator

Odixcity Consulting • Sud Africa

Remoto
ZAR 300.000 - 600.000
AI Trainer (Remote)
AI Trainer (Remote)

Hire Feed • Sud Africa

In loco
ZAR 350.000 - 550.000
Remote Business Analyst (Dutch) - 54883
Remote Business Analyst (Dutch) - 54883

Turing • Sud Africa

In loco
ZAR 30.000 - 50.000
Fully remote environment
Work on cutting-edge AI projects
Potential for contract extension based on performance
Senior AI LLM Engineer - Sandton - R1.2m PA
Senior AI LLM Engineer - Sandton - R1.2m PA

E-Merge • Sandton

In loco
ZAR 1.000.000 - 1.400.000
Computational Linguist Annotator
Computational Linguist Annotator

Odixcity Consulting • Sud Africa

Remoto
ZAR 350.000 - 550.000