LLM Evaluator (Model Response Analyst)

Odixcity Consulting

Portugal

Teletrabalho

EUR 40 000 - 60 000

Tempo integral

Há 7 dias
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Uma candidatura feita para esta oferta — um currículo e uma carta de apresentação personalizados que vão ao encontro do anúncio.

Ultrapassa os filtros ATS

Resumo da oferta

Odixcity Consulting seeks an LLM Evaluator to assess and improve large language models. You will judge factuality, coherence, safety, and alignment to guidelines, and provide clear feedback to training teams.

Responsibilities include ranking outputs, comparing prompts, and calibrating scoring with global evaluators to ensure reliability across the team.

Qualificações

  • 2+ years in NLP/AI QA, computational linguistics, data analysis or related field.
  • Experience crafting prompts to test model limits and behaviors.
  • Ability to explain why outputs are good or bad using logical criteria.

Responsabilidades

  • Evaluate and rank model-generated text using rubrics on factuality, coherence, safety, and alignment.
  • Review multiple outputs per prompt and justify preferred choices.
  • Provide actionable feedback to modeling teams on recurring failures.
  • Calibrate scoring with fellow evaluators in cross-check sessions.

Conhecimentos

NLP Evaluation
Prompt Engineering
Data Analysis
Quality Assurance
RLHF Data
Inter-rater Calibration

Formação académica

Bachelor's degree in Computer Science or related field

Descrição da oferta de emprego

Job Title: LLM Evaluator (Model Response Analyst)

Location: Remote (Worldwide)

Job Summary: We are seeking a detail-oriented and analytical LLM Evaluator to assess, analyze, and improve the performance of large language models (LLMs). In this role, you will evaluate AI-generated content for accuracy, coherence, factual reliability, bias, safety, and alignment with defined guidelines.

Responsibilities
  • Evaluate and rank model-generated text based based on complex rubrics covering dimensions such as factuality, coherence, safety, instruction- following, and creativity.
  • Review multiple model responses to the same prompt and determine which output a human would prefer, providing justifications for your choices.
  • Provide clear, concise feedback to the modeling and training teams regarding recurring failure models observed during evaluation sessions.
  • Attempt to “break” the model by crafting prompts designed to elicit biased, harmful, or insecure outputs to help patch safety vulnerabilities.
  • Collaborate with the quality assurance team to suggest improvements to evaluation guidelines when you encounter ambiguous or unclassifiable edge cases.
  • Participate in regular “cross-checking” sessions with other evaluators to calibrate scoring standards and ensure inter-rater reliability across the global team.
  • When a model underperforms, dig deeper than the surface score to hypothesize “why” the model made a specific error (e.g., training data vs. prompt misinterpretation).
  • Identify and flag novel or unexpected model behaviors to the research team, contributing to a living library of unique model outputs and failure modes.
Requirements
  • Minimum of 2 years of professional experience in a relevant field such as; Computational Linguistics, Data Analysis, Technical Writing, Quality Assurance (specifically for NLP/AI), or cognitive science.
  • Bachelor’s degree in Computer Science, or a relating field.
  • Deep understanding of how-to craft prompts to elicit specific behaviors and test model limits.
  • Ability to look at a text output and explain “why” it is “good” or “bad” based on logic, tone, factuality, and instruction adherence.
  • Experience working with Reinforcement Learning from Human Feedback (RLHF) data collection.
  • Proven experience monitoring and improving consistency among evaluation teams. Ability to analyze IAA scores and conduct calibration sessions to align judgement.
  • Experience sourcing, cleaning, and annotating datasets specifically for the fine-tuning or evaluating LLMs. Understanding of data distribution and its impact on model performance.
  • Familiarity with A/B testing concepts applied to AI. Ability to help design experiments to test if a new model version is truly “better” than the previous one.
Obtém a tua avaliação gratuita e confidencial do currículo.

ou arrasta e larga o ficheiro aqui.

Similar jobs

Ofertas semelhantes que vale a pena comparar

English Language Expert
English Language Expert

SME Careers • Portugal

Teletrabalho
EUR 34 000 - 62 000
English-Speaking AI Data Evaluation Specialist | Large Language Models
English-Speaking AI Data Evaluation Specialist | Large Language Models

Workster Jobs • Lisboa

Presencial
EUR 42 000 - 64 000
Private health insurance
Meal allowance
Professional development
Generalist Expert - Content Evaluator
Generalist Expert - Content Evaluator

Mercor • Lisboa

Presencial
EUR 50 000 - 65 000
AI Language QA Specialist — Remote Editor & Trainer
AI Language QA Specialist — Remote Editor & Trainer

SME Careers • Portugal

Teletrabalho
EUR 34 000 - 62 000
AI Content Evaluator: Detail-Led Generalist (UK/Europe)
AI Content Evaluator: Detail-Led Generalist (UK/Europe)

Mercor • Lisboa

Presencial
EUR 50 000 - 65 000
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Mercor • Lisboa

Presencial
EUR 65 000 - 100 000
Document Reviewer - Fully Remote
Document Reviewer - Fully Remote

Mercor • Lisboa

Teletrabalho
EUR 42 000 - 58 000
Senior Python AI Engineer (LLM & Multi-Agent Systems)
Senior Python AI Engineer (LLM & Multi-Agent Systems)

Seeking Alpha • Lisboa

Presencial
EUR 60 000 - 80 000
Flexible work environment
Remote work options
Variety of perks
Senior / Lead Machine Learning Engineer (AI & LLM Systems) - Hybrid (2 days office)
Senior / Lead Machine Learning Engineer (AI & LLM Systems) - Hybrid (2 days office)

HumanIT Digital Consulting • Lisboa

Presencial
EUR 70 000 - 100 000
Hybrid work model
Opportunity for technical leadership
Mentorship opportunities
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Mercor • Lisboa

Presencial
EUR 60 000 - 90 000