Senior Operations Specialist — AI Evaluation Pilot

Obsidian

Dallas (TX)

On-site

USD 55,000 - 110,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Mercor is running a short, structured evaluation pilot to measure how frontier AI models perform on realistic professional work — and how experienced practitioners judge that work. We are hiring 25 experts across four business domains. Each expert completes one task, then evaluates five model attempts at that same task.

This is a one-time study, not ongoing production work. No rubric and no golden answer are provided; professional judgment is the measure.

Qualifications

  • 5+ years of hands-on professional experience in Admin / Business Operations, Marketing, Human Resources, or Accounting.
  • Currently or recently practicing, so you can judge the work the way a working professional would.
  • Able to write a clear, specific, evidence-grounded rationale for a judgment.
  • Reliable within a short turnaround window.

Responsibilities

  • Sign an NDA before receiving any materials.
  • Solve one realistic, domain-specific task using the source files provided.
  • Review five model-generated outputs for that task. Rank them 1–5 with no ties, score each 0–100, and write evidence-based rationales that cite specific parts of the output.
  • Provide structured feedback on the task itself and on the study design.

Job description

Mercor is running a short, structured evaluation pilot to measure how frontier AI models perform on realistic professional work — and how experienced practitioners judge that work. We are hiring 25 experts across four business domains. Each expert completes one task, then evaluates five model attempts at that same task.

This is a one-time study, not ongoing production work. No rubric and no golden answer are provided; professional judgment is the measure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Business Operations Expert - Evaluator - AI Trainer
Business Operations Expert - Evaluator - AI Trainer

Obsidian • Dallas (TX)

On-site
USD 55,000 - 110,000
Data Science Expert - AI Evaluation
Data Science Expert - AI Evaluation

Mercor • New York (NY)

On-site
USD 140,000 - 180,000
Data Science Expert - Evaluation Specialist - AI Trainer
Data Science Expert - Evaluation Specialist - AI Trainer

Mercor • Philadelphia

On-site
USD 110,000 - 160,000
Data Science Expert - Evaluation Specialist
Data Science Expert - Evaluation Specialist

Mercor • New York (NY)

On-site
USD 120,000 - 160,000
AI-Powered Data Engineer & Model Evaluator
AI-Powered Data Engineer & Model Evaluator

Mercor • New York (NY)

On-site
USD 38,000 - 46,000
Data Science Expert - Evaluation Specialist
Data Science Expert - Evaluation Specialist

Obsidian • New York (NY)

On-site
USD 110,000 - 170,000
Data Science Expert - Evaluation Specialist - AI Trainer
Data Science Expert - Evaluation Specialist - AI Trainer

Obsidian • Philadelphia

On-site
USD 120,000 - 190,000
Remote AI Quality Evaluator, Support Operations
Remote AI Quality Evaluator, Support Operations

Mercor • New York (NY)

Remote
USD 34,000 - 55,000
AI Evaluation Scientist Math PhD Frontier Model Benchmark
AI Evaluation Scientist Math PhD Frontier Model Benchmark

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 179,000
Data Science Evaluation Specialist
Data Science Evaluation Specialist

Mercor • New York (NY)

On-site
USD 120,000 - 160,000