Generative AI Evaluator | $30/hr Remote

Crossing Hurdles

United States

Remote

CAD 27,552 - 41,328

Part time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

A remote-focused technology firm is looking for an individual proficient in evaluating large language model outputs. The role involves assessing AI systems, reviewing workflows, and providing actionable feedback to enhance product quality. Ideal candidates will have strong analytical skills, attention to detail, and excellent communication abilities in English. This position allows for flexibility, with a commitment between 10 to 40 hours each week, compensated hourly at $20–$30.

Qualifications

  • Strong experience in LLM evaluation, AI output analysis, QA/testing, UX research, or similar analytical roles.
  • Proficiency in rubric-based scoring, benchmarking frameworks, and AI quality assessment.
  • Excellent attention to detail with strong decision-making skills in ambiguous cases.
  • Proficient English communication skills (written and verbal).
  • Ability to work independently in a remote environment.
  • Comfortable committing to structured evaluation workflows and evolving guidelines.

Responsibilities

  • Evaluate outputs from large language models and autonomous agent systems using defined rubrics and quality standards.
  • Review multi-step agent workflows, including screenshots and reasoning traces, to assess accuracy and completeness.
  • Apply benchmarking criteria consistently while identifying edge cases and recurring failure patterns.
  • Provide structured, actionable feedback to support model refinement and product improvements.
  • Participate in calibration sessions to ensure consistent evaluation alignment across reviewers.
  • Adapt to evolving guidelines and ambiguous scenarios with sound judgment.
  • Document findings clearly and communicate insights to relevant stakeholders.

Skills

LLM evaluation
AI output analysis
QA/testing
UX research
Attention to detail
Proficient English communication

Job description

Type: Hourly contract

Compensation: $20–$30/hour

Location: Remote

Commitment: 10–40 hours/week

Role Responsibilities
  • Evaluate outputs from large language models and autonomous agent systems using defined rubrics and quality standards.
  • Review multi-step agent workflows, including screenshots and reasoning traces, to assess accuracy and completeness.
  • Apply benchmarking criteria consistently while identifying edge cases and recurring failure patterns.
  • Provide structured, actionable feedback to support model refinement and product improvements.
  • Participate in calibration sessions to ensure consistent evaluation alignment across reviewers.
  • Adapt to evolving guidelines and ambiguous scenarios with sound judgment.
  • Document findings clearly and communicate insights to relevant stakeholders.
Requirements
  • Strong experience in LLM evaluation, AI output analysis, QA/testing, UX research, or similar analytical roles.
  • Proficiency in rubric-based scoring, benchmarking frameworks, and AI quality assessment.
  • Excellent attention to detail with strong decision-making skills in ambiguous cases.
  • Proficient English communication skills (written and verbal).
  • Ability to work independently in a remote environment.
  • Comfortable committing to structured evaluation workflows and evolving guidelines.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote | AI Evaluation Specialist — $30–$90/hour
Remote | AI Evaluation Specialist — $30–$90/hour

24-Mag Llc • Northern (KY)

Hybrid
USD 41,000 - 124,000
Generative AI Quality Evaluator (Remote)
Generative AI Quality Evaluator (Remote)

Crossing Hurdles • United States

Remote
CAD 60,000 - 80,000
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

Remote
Secure computer and high-speed internet required
AI Evaluation Specialist
AI Evaluation Specialist

micro1 • United States

Remote
AUD 70,000 - 110,000
Remote AI Quality Evaluator & Feedback Specialist
Remote AI Quality Evaluator & Feedback Specialist

24-Mag Llc • Northern (KY)

Hybrid
USD 41,000 - 124,000
Remote AI Training Lead: Quality & Evaluation
Remote AI Training Lead: Quality & Evaluation

YO IT Consulting • New York (NY)

On-site
USD 55,104 - 96,432
Remote | Technical Writer / Editor — $90–$140/hour
Remote | Technical Writer / Editor — $90–$140/hour

24-Mag Llc • Northern (KY)

Hybrid
USD 124,000 - 193,000
Remote contractor
Open to US, Canada, UK
Remote | Product Manager/Product Owner — $90–$140/hour
Remote | Product Manager/Product Owner — $90–$140/hour

24-Mag Llc • New York (NY), Northern (KY)

Hybrid
USD 110,000 - 179,000
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Mississippi

Remote
USD 55,000 - 75,000
Competitive pay up to $80/hr
Flexible, remote work
Participation in advanced AI projects
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Austin (TX)

Remote
Competitive pay
Flexible schedule
Experience in advanced AI projects