Stand out for this role — generate a tailored resume and cover letter in about a minute.
Mercor is conducting a short, structured evaluation pilot to measure how frontier AI models perform on realistic professional work. We are hiring 25 experts across four domains, each expert completes one task and evaluates five model attempts for that same task.
This is a one-time study, not ongoing production work. We seek professionals with 5+ years of hands-on experience in Admin/Business Operations, Marketing, HR, or Accounting, who can judge like a working professional and write clear,
Mercor is running a short, structured evaluation pilot to measure how frontier AI models perform on realistic professional work — and how experienced practitioners judge that work. We are hiring 25 experts across four business domains. Each expert completes one task, then evaluates five model attempts at that same task. This is a one-time study, not ongoing production work.
Sign an NDA before receiving any materials.
Solve one realistic, domain-specific task using the source files provided.
Review five model-generated outputs for that task. Rank them 1–5 with no ties, score each 0–100, and write evidence-based rationales that cite specific parts of the output.
Provide structured feedback on the task itself and on the study design.
Up to 13 hours total: 3–10 hours to solve the task, approximately 2 hours to rank and score the five outputs, and approximately 1 hour to provide feedback.
No rubric and no golden answer are provided. Your professional judgment is the measurement.
Use of LLMs or other AI assistants is prohibited at every stage — solving, ranking, scoring, and writing rationales.
Admin / Business Operations
Marketing
Human Resources
Accounting
5+ years of hands-on professional experience in one of the four domains above.
Currently or recently practicing, so you can judge the work the way a working professional would.
Able to write a clear, specific, evidence-grounded rationale for a judgment.
Reliable within a short turnaround window.
You are not eligible for this pilot if you have worked on Project Alchemy in any capacity — task author, reviewer, world expert, or anyone who has had access to its source world data. Prior exposure to that material would invalidate the study.