Business Operations Expert - Evaluator - AI Trainer

Obsidian

Dallas (TX)

On-site

USD 55,000 - 110,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Mercor is running a short, structured evaluation pilot to measure how frontier AI models perform on realistic professional work — and how experienced practitioners judge that work. We are hiring 25 experts across four business domains. Each expert completes one task, then evaluates five model attempts at that same task.

This is a one-time study, not ongoing production work. No rubric and no golden answer are provided; professional judgment is the measure.

Qualifications

  • 5+ years of hands-on professional experience in Admin / Business Operations, Marketing, Human Resources, or Accounting.
  • Currently or recently practicing, so you can judge the work the way a working professional would.
  • Able to write a clear, specific, evidence-grounded rationale for a judgment.
  • Reliable within a short turnaround window.

Responsibilities

  • Sign an NDA before receiving any materials.
  • Solve one realistic, domain-specific task using the source files provided.
  • Review five model-generated outputs for that task. Rank them 1–5 with no ties, score each 0–100, and write evidence-based rationales that cite specific parts of the output.
  • Provide structured feedback on the task itself and on the study design.

Job description

Role Overview

Mercor is running a short, structured evaluation pilot to measure how frontier AI models perform on realistic professional work — and how experienced practitioners judge that work. We are hiring 25 experts across four business domains. Each expert completes one task, then evaluates five model attempts at that same task. This is a one-time study, not ongoing production work.

What You'll Do
  • Sign an NDA before receiving any materials.
  • Solve one realistic, domain-specific task using the source files provided.
  • Review five model-generated outputs for that task. Rank them 1–5 with no ties, score each 0–100, and write evidence-based rationales that cite specific parts of the output.
  • Provide structured feedback on the task itself and on the study design.
Time Commitment

Up to 13 hours total: 3–10 hours to solve the task, approximately 2 hours to rank and score the five outputs, and approximately 1 hour to provide feedback.

Study Conditions
  • No rubric and no golden answer are provided. Your professional judgment is the measurement.
  • Use of LLMs or other AI assistants is prohibited at every stage — solving, ranking, scoring, and writing rationales.
Domains we are focused on
  • Admin / Business Operations
  • Marketing
  • Human Resources
  • Accounting
Who We're Looking For
  • 5+ years of hands‑on professional experience in one of the four domains above.
  • Currently or recently practicing, so you can judge the work the way a working professional would.
  • Able to write a clear, specific, evidence‑grounded rationale for a judgment.
  • Reliable within a short turnaround window.
Eligibility Restriction

You are not eligible for this pilot if you have worked on Project Alchemy in any capacity — task author, reviewer, world expert, or anyone who has had access to its source world data. Prior exposure to that material would invalidate the study.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Operations Specialist — AI Evaluation Pilot
Senior Operations Specialist — AI Evaluation Pilot

Obsidian • Dallas (TX)

On-site
USD 55,000 - 110,000
Data Science Expert - Evaluation Specialist - AI Trainer
Data Science Expert - Evaluation Specialist - AI Trainer

Mercor • Philadelphia

On-site
USD 110,000 - 160,000
Data Science Expert - Evaluation Specialist
Data Science Expert - Evaluation Specialist

Mercor • New York (NY)

On-site
USD 120,000 - 160,000
Data Science Expert - Evaluation Specialist - AI Trainer
Data Science Expert - Evaluation Specialist - AI Trainer

Obsidian • Philadelphia

On-site
USD 120,000 - 190,000
Data Science Expert - AI Evaluation
Data Science Expert - AI Evaluation

Mercor • New York (NY)

On-site
USD 140,000 - 180,000
Sales Engineering Expert - Evaluator
Sales Engineering Expert - Evaluator

Mercor • New York (NY)

On-site
USD 120,000 - 180,000
Marketing Expert - Creator Programs - AI Trainer
Marketing Expert - Creator Programs - AI Trainer

Mercor • Houston (TX)

On-site
USD 70,000 - 110,000
Marketing Expert - Creator Programs - AI Trainer
Marketing Expert - Creator Programs - AI Trainer

Obsidian • Houston (TX)

On-site
USD 90,000 - 130,000
Sales Engineering Expert - Evaluator
Sales Engineering Expert - Evaluator

Obsidian • New York (NY)

On-site
USD 100,000 - 140,000
Accounting Specialist - Fully Remote
Accounting Specialist - Fully Remote

Mercor • United States

Remote