AI Agent Workflow Evaluator

OpenTrain AI

Northern (KY)

Hybrid

USD 41,000 - 124,000

Part time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

OpenTrain AI is seeking an AI Agent Workflow Evaluator for remote, part-time work. You will test AI assistants (ChatGPT, Claude) in complex business workflows, scoring results with rubrics and recording detailed feedback.

The role focuses on clear written communication and accurate documentation of processes. Requires at least five years in a technology-enabled business function, a bachelor’s degree, and strong English writing.

Qualifications

  • Minimum five years of professional experience in a business function where technology is used to solve operational challenges.
  • Daily, hands-on professional use of ChatGPT, Claude, or both
  • Experience creating, applying, or reviewing evaluation rubrics, QA scorecards, grading criteria, or content review guidelines
  • Excellent written English and ability to document findings precisely and provide clear, actionable feedback

Responsibilities

  • Reproduce authentic workplace use cases by connecting AI assistants with business tools
  • Maintain accurate records so evaluations are transparent, consistent, and reproducible
  • Score generated outputs using rubrics, QA scorecards, or grading criteria
  • Write precise, actionable feedback about model performance
  • Identify recurring strengths, weaknesses, and opportunities across workflows
  • Connect AI assistants with tools like Google Drive, Gmail, Slack, and Notion
  • Record each workflow step and compare AI behavior with operational requirements

Skills

Excellent written English
Document findings precisely
Using ChatGPT/Claude
Create/evaluate rubrics
QA scorecards

Education

Bachelor's degree or higher

Tools

ChatGPT
Claude
Google Drive
Gmail
Slack
Notion

Job description

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover projects, build a professional profile, and apply in minutes.

Creating an OpenTrain account is free. Your profile can help you showcase relevant AI training experience and grow a career in a fast-moving field where human judgment directly improves how AI systems work.

About AI Training Work

AI training is the human side of building artificial intelligence. People evaluate model responses, write feedback, and judge whether AI outputs are accurate, useful, complete, and relevant. This work helps shape the behavior and reliability of modern AI systems.

This opportunity focuses on evaluating generative AI assistants in realistic professional settings. It is remote and offers flexible contractor work for contributors who can commit at least 20 hours per week.

The Role

As an AI Agent Workflow Evaluator, you will test AI assistants such as ChatGPT and Claude through complex, multi-step business workflows. You will assess how well each assistant handles practical operational requirements and document the results clearly.

Your evaluations will help identify strengths, weaknesses, and opportunities to improve the quality, completeness, relevance, and reliability of next-generation AI systems. Previous AI training experience is not required.

  • Employment type: Remote contractor and part-time
  • Time commitment: 20+ hours per week
  • Compensation: $30-$90 per hour
  • Work authorization: United States, Canada, United Kingdom, Ireland, Australia, or New Zealand
  • United States strongly preferred
  • Working language: English
What You'll Do

You will reproduce authentic workplace use cases by connecting AI assistants with business and productivity tools. You will maintain accurate records so that your evaluations are transparent, consistent, and reproducible.

Using defined rubrics and evaluation criteria, you will score AI-generated outputs and provide constructive feedback that points to specific improvements. You will also track recurring patterns in assistant behavior across different workflows.

  • Run complex, multi-step business scenarios that reflect genuine professional workflows
  • Use ChatGPT, Claude, or both to complete realistic tasks
  • Record each step of a workflow and compare AI behavior with operational requirements
  • Score generated outputs using rubrics, QA scorecards, grading criteria, or review guidelines
  • Write precise, actionable feedback about model performance
  • Identify recurring strengths, weaknesses, and process opportunities
  • Connect AI assistants with tools such as Google Drive, Gmail, Slack, and Notion
  • Maintain detailed written documentation that supports transparency and reproducibility
Requirements

This role is suited to professionals who understand how technology supports real business operations and who can assess work against clear quality standards. You should be comfortable investigating multi-step processes, making careful judgments, and explaining your reasoning in written English.

  • At least five years of professional experience in a business function where technology is used to solve operational challenges
  • Completed bachelor's degree or higher in any discipline
  • Daily, hands-on professional use of ChatGPT, Claude, or both
  • Experience creating, applying, or reviewing evaluation rubrics, QA scorecards, grading criteria, or content review guidelines
  • Comfort connecting AI assistants with workplace and productivity software
  • Excellent written English
  • Ability to document findings precisely and provide clear, actionable feedback
Helpful Background

Experience in any of the following areas may help you contribute effectively. These backgrounds can provide useful practice in applying consistent standards, reviewing outputs, or documenting complex processes.

  • Academic grading
  • Quality assurance
  • Hiring scorecards
  • Content moderation
  • Annotation guidelines
  • AI model evaluation
  • Documenting multi-step professional processes
  • Judging outputs for quality, completeness, and relevance
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Personalized AI Assistant Evaluation Expert
Personalized AI Assistant Evaluation Expert

OpenTrain AI • Northern (KY)

Hybrid
USD 69,000 - 276,000
Remote contract
Flexible schedule
20–40 hours per week
Remote AI Workflow Evaluator, Part-Time
Remote AI Workflow Evaluator, Part-Time

OpenTrain AI • Northern (KY)

Hybrid
USD 41,000 - 124,000
AI Analytics Workflow Evaluator
AI Analytics Workflow Evaluator

OpenTrain AI • Snowflake (AZ), Northern (KY)

Hybrid
USD 69,000 - 83,000
Remote work worldwide
Part-time contractor
20+ hours per week
+1
AI Software Development Trace Evaluator
AI Software Development Trace Evaluator

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000
Personalized AI Response Evaluator
Personalized AI Response Evaluator

OpenTrain AI • Northern (KY)

Hybrid
USD 21,000 - 34,000
AI Evaluation Specialist
AI Evaluation Specialist

micro1 • United States

Remote
AUD 70,000 - 110,000
Public Sector AI Work Product Reviewer
Public Sector AI Work Product Reviewer

OpenTrain AI • Northern (KY)

Hybrid
USD 69,000 - 76,000
Remote work
Flexible hours
Part-time contractor
Personalized AI Response Evaluation Rater
Personalized AI Response Evaluation Rater

OpenTrain AI • Northern (KY)

Hybrid
USD 17,000 - 28,000
GTM and Revenue Operations Document Evaluator
GTM and Revenue Operations Document Evaluator

OpenTrain AI • Northern (KY)

Hybrid
USD 110,000 - 220,000
Remote work
40 hours per week baseline
Flexible schedule
Product Manager / Product Owner AI Evaluator
Product Manager / Product Owner AI Evaluator

OpenTrain AI • Northern (KY)

Hybrid
USD 124,000 - 193,000