AI Agent Evaluator - Remote Agent Run Review

AuraOne

United States

Remote

USD 28,000 - 55,000

Full time

30 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

AuraOne is seeking an AI Agent Evaluator to work remotely as a Contractor. You will review AI agent outputs, apply a versioned rubric, and provide structured feedback, labels, and rationales to help retrain models.

Responsibilities include comparing paired responses, tagging content issues, and documenting recurring failure modes for calibration. The role is remote and US-eligible with an hourly rate paid after interview.

Qualifications

  • Experience in evaluating AI agent outputs using structured rubrics.
  • Ability to write clear rationale names for rubric clauses.
  • Attention to detail and the ability to flag prompt issues.

Responsibilities

  • Evaluate AI agent outputs against a rubric and assign severity tags.
  • Compare paired responses and justify the stronger answer with rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with tags.
  • Document prompts and rubric gaps for rubric updates.
  • Calibrate reviewer quality against gold-standard examples weekly.
  • Record recurring failure modes for model retraining.

Skills

Model output evaluation
Rubric-based annotation
Severity tagging
Inter-rater calibration
AI Agent evaluation
AI model evaluation

Job description

Category: Frontier Model Evaluation · Pay: Hourly rate confirmed after the interview process · Location: Remote — US-eligible · Contractor

AI Agent Evaluator - Remote Agent Run Review is a remote evaluation track for reviewing ai agent evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.

Why this role matters

AI data reviewers help turn ai agent evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.

Responsibilities
  • Evaluate ai agent evaluation model outputs against a versioned rubric and assign severity tags for AI Agent Evaluator - Remote Agent Run Review assignments.
  • Compare paired responses and pick the stronger answer with a written rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with structured tags.
  • Capture ambiguous prompts and route them back to the program team for rubric updates.
  • Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
  • Document recurring failure modes so the modeling team can target them in the next training run.
Qualifications
  • Prior evaluation, annotation, or human-rater experience on ai agent evaluation or adjacent content for AI Agent Evaluator - Remote Agent Run Review work.
  • Comfort applying multi-page rubrics consistently across long batches.
  • Clear written reasoning that names the issue and the rubric clause being applied.
  • Strong attention to detail and the ability to flag when a prompt itself is the problem.
  • Reliable async availability for at least 10 hours per week.
Example tasks
  • Compare two ai agent evaluation model responses to the same prompt and pick the stronger one with rationale.
  • Tag an unsafe response with the correct policy category and severity.
  • Audit a 50-row batch for rubric consistency and report drift to the program lead.
  • Propose a rubric clarification after spotting a recurring failure mode.
Nice to have
  • Background in linguistics, content moderation, or trust & safety review.
  • Experience with inter-rater agreement metrics and calibration cycles.
  • Domain expertise that lets you spot subject-matter errors automated checks miss.
Skills
  • Model output evaluation
  • Rubric-based annotation
  • Severity tagging
  • Inter-rater calibration
  • AI Agent evaluation
  • AI model evaluation
Work model

Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.

Compensation

Hourly rate confirmed after the interview process.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Trainer & Evaluator
AI Trainer & Evaluator

AuraOne • United States

Remote
USD 34,000 - 83,000
Remote work
Contractor engagement
US-eligible
CRM Agent Workflow AI Reviewer
CRM Agent Workflow AI Reviewer

AuraOne, Inc. • United States

Remote
USD 2,755,000 - 4,133,000
Agent Workflow Evaluator
Agent Workflow Evaluator

AuraOne • United States

On-site
USD 28,000 - 55,000
AI Agent Interaction Specialist - Dialogue & Conversation Evalua
AI Agent Interaction Specialist - Dialogue & Conversation Evalua

AuraOne • United States

On-site
USD 28,000 - 50,000
AI Trainer
AI Trainer

AuraOne • United States

On-site
USD 28,000 - 55,000
Remote AI Agent Evaluator & Rubric Specialist
Remote AI Agent Evaluator & Rubric Specialist

AuraOne • United States

Remote
USD 28,000 - 55,000
Computer-Use Agent Evaluator
Computer-Use Agent Evaluator

AuraOne • United States

On-site
USD 28,000 - 41,000
Remote work
Contractor arrangement
AI Data Annotation Expert
AI Data Annotation Expert

AuraOne • United States

On-site
USD 34,000 - 62,000
Conversational AI Evaluator
Conversational AI Evaluator

AuraOne • United States

On-site
USD 34,000 - 55,000
Human Resources Operations AI Evaluator
Human Resources Operations AI Evaluator

AuraOne • United States

On-site
USD 28,000 - 55,000