Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
OpenTrain AI is seeking an AI Agent Workflow Evaluator for remote, part-time work. You will test AI assistants (ChatGPT, Claude) in complex business workflows, scoring results with rubrics and recording detailed feedback.
The role focuses on clear written communication and accurate documentation of processes. Requires at least five years in a technology-enabled business function, a bachelor’s degree, and strong English writing.
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover projects, build a professional profile, and apply in minutes.
Creating an OpenTrain account is free. Your profile can help you showcase relevant AI training experience and grow a career in a fast-moving field where human judgment directly improves how AI systems work.
AI training is the human side of building artificial intelligence. People evaluate model responses, write feedback, and judge whether AI outputs are accurate, useful, complete, and relevant. This work helps shape the behavior and reliability of modern AI systems.
This opportunity focuses on evaluating generative AI assistants in realistic professional settings. It is remote and offers flexible contractor work for contributors who can commit at least 20 hours per week.
As an AI Agent Workflow Evaluator, you will test AI assistants such as ChatGPT and Claude through complex, multi-step business workflows. You will assess how well each assistant handles practical operational requirements and document the results clearly.
Your evaluations will help identify strengths, weaknesses, and opportunities to improve the quality, completeness, relevance, and reliability of next-generation AI systems. Previous AI training experience is not required.
You will reproduce authentic workplace use cases by connecting AI assistants with business and productivity tools. You will maintain accurate records so that your evaluations are transparent, consistent, and reproducible.
Using defined rubrics and evaluation criteria, you will score AI-generated outputs and provide constructive feedback that points to specific improvements. You will also track recurring patterns in assistant behavior across different workflows.
This role is suited to professionals who understand how technology supports real business operations and who can assess work against clear quality standards. You should be comfortable investigating multi-step processes, making careful judgments, and explaining your reasoning in written English.
Experience in any of the following areas may help you contribute effectively. These backgrounds can provide useful practice in applying consistent standards, reviewing outputs, or documenting complex processes.