Remote Agent Workflow Evaluator & Rubric Auditor

AI Trainer Jobs

United States

Remote

USD 28,000 - 55,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

AI Trainer Jobs is seeking an Agent Workflow Evaluator for a remote evaluation track. You will review agent workflow evaluation prompts and responses against a quality rubric, compare paired outputs, and provide structured feedback to help retrain models.

Responsibilities include labeling edge cases, flagging hallucinations, and documenting rationales and regression cases. This is a contractor role with flexible hours and compensation arranged after interview.

Qualifications

  • Prior evaluation, annotation, or human-rater experience on agent workflow evaluation or adjacent content.
  • Comfort applying multi-page rubrics consistently across long batches.
  • Clear written reasoning that names the issue and the rubric clause being applied.
  • Strong attention to detail and the ability to flag when a prompt itself is the problem.
  • Reliable async availability for at least 10 hours per week.

Responsibilities

  • Evaluate agent workflow evaluation model outputs against a versioned rubric and assign severity tags.
  • Compare paired responses and pick the stronger answer with written rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with structured tags.
  • Capture ambiguous prompts and route them back to the program team for rubric updates.
  • Maintain reviewer-quality scores by calibrating against gold-standard examples each week.

Skills

Evaluation experience
Annotation experience
Written reasoning
Attention to detail
Async availability
Remote work

Job description

AI Trainer Jobs is seeking an Agent Workflow Evaluator for a remote evaluation track. You will review agent workflow evaluation prompts and responses against a quality rubric, compare paired outputs, and provide structured feedback to help retrain models.

Responsibilities include labeling edge cases, flagging hallucinations, and documenting rationales and regression cases. This is a contractor role with flexible hours and compensation arranged after interview.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Agent Workflow Evaluator
Agent Workflow Evaluator

AI Trainer Jobs • United States

Remote
USD 28,000 - 55,000
Remote AI Evaluation Specialist - Rubric & Labeling
Remote AI Evaluation Specialist - Rubric & Labeling

AI Trainer Jobs • United States

Remote
USD 28,000 - 41,000
Remote work
Contractor arrangement
Remote HR Operations AI Evaluator & Rubric Reviewer
Remote HR Operations AI Evaluator & Rubric Reviewer

AI Trainer Jobs • United States

Remote
USD 28,000 - 55,000
Remote Data Evaluator for Model Feedback & Rubrics
Remote Data Evaluator for Model Feedback & Rubrics

AI Trainer Jobs • United States

Remote
USD 21,000 - 34,000
Remote AI Workflow Evaluator - Management Consultant
Remote AI Workflow Evaluator - Management Consultant

AI Trainer Jobs • United States

Remote
USD 41,000 - 76,000
Remote Fact-Check Workflow Evaluator (Rubric-Based)
Remote Fact-Check Workflow Evaluator (Rubric-Based)

AI Trainer Jobs • United States

Remote
USD 28,000 - 55,000
Remote Drone Ops AI Evaluator & Rubric Auditor
Remote Drone Ops AI Evaluator & Rubric Auditor

AI Trainer Jobs • United States

Remote
USD 28,000 - 55,000
Remote Conversational AI Evaluator & Rubric Labeler
Remote Conversational AI Evaluator & Rubric Labeler

AI Trainer Jobs • United States

Remote
USD 34,000 - 55,000
Remote HRI Evaluator: Quality Metrics & Rubric Calibration
Remote HRI Evaluator: Quality Metrics & Rubric Calibration

AI Trainer Jobs • United States

Remote
USD 34,000 - 55,000
Remote Agent Regression QA & Evaluation Specialist
Remote Agent Regression QA & Evaluation Specialist

AuraOne • United States

On-site
USD 34,440 - 82,656