Remote RLHF Data Reviewer: Consensus & Rubric Feedback

AI Trainer Jobs

United States

Remote

USD 28,000 - 55,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

AuraOne is seeking a Labeler Consensus Preference Data Reviewer for a remote evaluation track. You will review labeler consensus preference data evaluation prompts and responses against our quality rubric, compare paired outputs, and write structured feedback the modeling team can use to retrain.

You will produce preference rankings, reward-model feedback, and calibrated human judgment for post-training pipelines, tagging edge cases and risk signals, while ensuring consistency across long

Qualifications

  • Prior evaluation, annotation, or human-rater experience on labeler consensus data evaluation or adjacent content.
  • Ability to apply multi-page rubrics consistently across long batches.
  • Clear written reasoning that names the issue and the rubric clause being applied.

Responsibilities

  • Evaluate labeler consensus preference data evaluation model outputs against a versioned rubric and assign severity tags for Labeler Consensus Preference Data Reviewer assignments.
  • Compare paired responses and pick the stronger answer with a written rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with structured tags.
  • Capture ambiguous prompts and route them back to the program team for rubric updates.

Job description

AuraOne is seeking a Labeler Consensus Preference Data Reviewer for a remote evaluation track. You will review labeler consensus preference data evaluation prompts and responses against our quality rubric, compare paired outputs, and write structured feedback the modeling team can use to retrain.

You will produce preference rankings, reward-model feedback, and calibrated human judgment for post-training pipelines, tagging edge cases and risk signals, while ensuring consistency across long

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Labeler Consensus Preference Data Reviewer
Labeler Consensus Preference Data Reviewer

AI Trainer Jobs • United States

Remote
USD 28,000 - 55,000
Remote RLHF Evaluator: Preference Taxonomy Feedback
Remote RLHF Evaluator: Preference Taxonomy Feedback

AI Trainer Jobs • United States

Remote
USD 34,000 - 55,000
Senior RLHF Preference Data Reviewer - Remote
Senior RLHF Preference Data Reviewer - Remote

AI Trainer Jobs • United States

Remote
USD 34,000 - 55,000
Remote RLHF QA Evaluator - Preference Data
Remote RLHF QA Evaluator - Preference Data

AI Trainer Jobs • United States

Remote
USD 28,000 - 55,000
Remote work
Remote RLHF: Reward Hacking Preference Data Reviewer
Remote RLHF: Reward Hacking Preference Data Reviewer

AI Trainer Jobs • United States

Remote
USD 27,552,000 - 41,328,000
Remote work (US-eligible)
Contractor role
Remote Video Data Labeler – Content & Evaluation
Remote Video Data Labeler – Content & Evaluation

AI Trainer Jobs • United States

Remote
USD 21,000 - 48,000
Remote Label Quality Ontology QA Reviewer
Remote Label Quality Ontology QA Reviewer

AI Trainer Jobs • United States

Remote
USD 110,000 - 165,000
Remote AI Preference Data QA Reviewer & Feedback Specialist
Remote AI Preference Data QA Reviewer & Feedback Specialist

AI Trainer Jobs • United States

Remote
USD 28,000 - 48,000
Remote AI Data Evaluation & Annotation Expert
Remote AI Data Evaluation & Annotation Expert

AI Trainer Jobs • United States

Remote
USD 34,000 - 62,000
Label Quality Ontology QA Reviewer
Label Quality Ontology QA Reviewer

AI Trainer Jobs • United States

Remote
USD 110,000 - 165,000