Remote Pairwise Preference Evaluator & Feedback Specialist

AuraOne

United States

On-site

USD 27,552 - 55,104

Part time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

AuraOne is seeking a remote, contractor-based Pairwise Preference Reward Model Evaluator to review prompts and responses against the company's quality rubric. You will compare paired outputs, label edge cases, and draft structured feedback to help retrain the modeling team.

The role emphasizes careful reasoning, consistency across long review batches, and clear documentation of errors. You will work asynchronously for at least 10 hours weekly, with compensation discussed after the interview.

Qualifications

  • Prior evaluation, annotation, or human-rater experience on pairwise reward model evaluation or adjacent content.
  • Comfort applying multi-page rubrics consistently across long batches.
  • Clear written reasoning that names the issue and the rubric clause being applied.
  • Strong attention to detail and the ability to flag when a prompt itself is the problem.
  • Reliable async availability for at least 10 hours per week.

Responsibilities

  • Evaluate pairwise preference reward model evaluation model outputs against a versioned rubric and assign severity tags for Pairwise Preference Reward Model Evaluator assignments.
  • Compare paired responses and pick the stronger answer with a written rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with structured tags.
  • Capture ambiguous prompts and route them back to the program team for rubric updates.
  • Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
  • Document recurring failure modes so the modeling team can target them in the next training run.

Skills

Model output evaluation
Rubric-based annotation
Severity tagging
Inter-rater calibration
Pairwise evaluation
RLHF
Rater calibration

Job description

AuraOne is seeking a remote, contractor-based Pairwise Preference Reward Model Evaluator to review prompts and responses against the company's quality rubric. You will compare paired outputs, label edge cases, and draft structured feedback to help retrain the modeling team.

The role emphasizes careful reasoning, consistency across long review batches, and clear documentation of errors. You will work asynchronously for at least 10 hours weekly, with compensation discussed after the interview.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Evaluator: Preference Taxonomy & Model Quality
Remote Evaluator: Preference Taxonomy & Model Quality

AuraOne • United States

On-site
USD 34,440 - 55,104
Remote Formal Spec Model Evaluator (Contractor)
Remote Formal Spec Model Evaluator (Contractor)

AuraOne • United States

On-site
USD 27,552 - 55,104
Remote AI Evaluation & Feedback Specialist
Remote AI Evaluation & Feedback Specialist

AuraOne • United States

On-site
USD 55,000 - 110,000
Remote work
US-eligible
Contractor position
Remote Preference Dataset QA Evaluator RLHF
Remote Preference Dataset QA Evaluator RLHF

AuraOne • United States

On-site
USD 55,104 - 82,656
Remote English Language Evaluation Specialist: Feedback
Remote English Language Evaluation Specialist: Feedback

AuraOne • United States

On-site
USD 34,440 - 55,104
Remote Mobile AI Evaluation Specialist
Remote Mobile AI Evaluation Specialist

AuraOne • United States

On-site
USD 34,000 - 69,000
Remote Retrieval Relevance Evaluator — Quality Feedback
Remote Retrieval Relevance Evaluator — Quality Feedback

AuraOne • United States

On-site
USD 34,440 - 55,104
Remote Human Data Evaluator & Rubric Reviewer
Remote Human Data Evaluator & Rubric Reviewer

AuraOne • United States

On-site
USD 34,000 - 55,000
Remote Prosody Evaluation Specialist – Model Output QA
Remote Prosody Evaluation Specialist – Model Output QA

AuraOne • United States

On-site
USD 27,552 - 55,104
Remote AI Preference Data QA Specialist
Remote AI Preference Data QA Specialist

AuraOne • United States

On-site
USD 34,440 - 55,104