Turn this role into an interview — a resume and cover letter built around what this employer wants.
Mercor is running a short, structured evaluation pilot to measure how frontier AI models perform on realistic professional work — and how experienced practitioners judge that work. We are hiring 25 experts across four business domains. Each expert completes one task, then evaluates five model attempts at that same task.
This is a one-time study, not ongoing production work. No rubric and no golden answer are provided; professional judgment is the measure.
Mercor is running a short, structured evaluation pilot to measure how frontier AI models perform on realistic professional work — and how experienced practitioners judge that work. We are hiring 25 experts across four business domains. Each expert completes one task, then evaluates five model attempts at that same task. This is a one-time study, not ongoing production work.
Up to 13 hours total: 3–10 hours to solve the task, approximately 2 hours to rank and score the five outputs, and approximately 1 hour to provide feedback.
You are not eligible for this pilot if you have worked on Project Alchemy in any capacity — task author, reviewer, world expert, or anyone who has had access to its source world data. Prior exposure to that material would invalidate the study.