Tool-Use & AI Evaluation Specialist (Remote)

AI Trainer Jobs

United States

Remote

USD 28,000 - 55,000

Part time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Remote work
Hourly contractor engagement

Job summary

AuraOne is seeking remote reviewers to evaluate tool use and computer use prompts and responses against our quality rubric. As a remote contractor, you will review frontier model outputs, label edge cases, and provide structured feedback to retrain the model.

Ideal candidates have prior evaluation or annotation experience, strong attention to detail, and the ability to work asynchronously for at least 10 hours per week; compensation is hourly after interview.

Qualifications

  • Prior evaluation, annotation, or human-rater experience in tool use and computer use evaluation.
  • Ability to apply multi-page rubrics consistently across long batches.
  • Clear written reasoning naming the issue and rubric clause.
  • Strong attention to detail and the ability to flag issues.
  • Reliable async availability for at least 10 hours per week.

Responsibilities

  • Evaluate tool use and computer use evaluation model outputs against rubric and assign severity tags for Tool-Use and Computer-Use Evaluator assignments.
  • Compare paired responses and pick the stronger answer with a written rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with structured tags.
  • Capture ambiguous prompts and route them back to the program team for rubric updates.
  • Maintain reviewer-quality scores by calibrating against gold-standard examples each week.

Skills

Tool use evaluation
Annotation experience
Inter-rater calibration
Written reasoning
Async availability

Job description

AuraOne is seeking remote reviewers to evaluate tool use and computer use prompts and responses against our quality rubric. As a remote contractor, you will review frontier model outputs, label edge cases, and provide structured feedback to retrain the model.

Ideal candidates have prior evaluation or annotation experience, strong attention to detail, and the ability to work asynchronously for at least 10 hours per week; compensation is hourly after interview.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Specialist — Frontiers QA (Remote)
AI Evaluation Specialist — Frontiers QA (Remote)

AI Trainer Jobs • United States

Remote
USD 34,000 - 49,000
Remote Visual Evaluation & Rubric Feedback Specialist
Remote Visual Evaluation & Rubric Feedback Specialist

AI Trainer Jobs • United States

Remote
USD 27,552,000 - 49,594,000
Remote Senior AI Evaluation & Feedback Specialist
Remote Senior AI Evaluation & Feedback Specialist

AuraOne • United States

On-site
USD 34,000 - 83,000
Remote Human Data QA & Evaluation Reviewer
Remote Human Data QA & Evaluation Reviewer

AI Trainer Jobs • United States

Remote
USD 34,000 - 55,000
Remote work
Contractor role
Flexible hours
Remote AI Data Reviewer - Windows Generalist Evaluations
Remote AI Data Reviewer - Windows Generalist Evaluations

AuraOne • United States

On-site
USD 41,000 - 55,000
Senior AI Evaluation Specialist — Remote
Senior AI Evaluation Specialist — Remote

AI Trainer Jobs • United States

Remote
USD 34,000 - 55,000
Remote Technical Writing Evaluation Specialist
Remote Technical Writing Evaluation Specialist

AuraOne • United States

On-site
USD 34,000 - 55,000
Remote Video AI Evaluation Specialist
Remote Video AI Evaluation Specialist

AI Trainer Jobs • United States

Remote
USD 34,000 - 55,000
Remote Computer Vision Evaluation Specialist
Remote Computer Vision Evaluation Specialist

AI Trainer Jobs • United States

Remote
USD 28,000 - 55,000
Remote IT Helpdesk AI Evaluation Auditor
Remote IT Helpdesk AI Evaluation Auditor

AI Trainer Jobs • United States

Remote
USD 34,000 - 62,000