Mathematics PhD - Benchmark Specialist

Mercor

New York (NY)

On-site

USD 83,000 - 179,000

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote work
Flexible hours
Collaborative AI research environment

Job summary

Mercor is seeking expert mathematicians to author and review high-quality academic assessment content for an AI research initiative. You will author or verify rigorous multiple-choice questions across core mathematics domains, evaluate solution quality, and help establish gold-standard benchmarks for advancing AI capabilities.

You will require deep mathematical understanding, strong written English, and the ability to produce precise problem statements, solutions, and references.

Qualifications

  • PhD or doctoral candidate in Mathematics, Applied Mathematics, Statistics, or closely related field.
  • Master's degree considered for exceptional depth.
  • Excellent written English and ability to express complex ideas clearly.

Responsibilities

  • Author original math questions that test deep conceptual understanding.
  • Rate each question's difficulty: Medium, Hard, or Expert.
  • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives.
  • Write step-by-step Chain-of-Thought solutions in Markdown format.
  • Supply 1–5 academic references per question.
  • Flag issues with clarity, completeness, precision, or solvability and justify edits.

Skills

Graduate math concepts
Formal proof writing
Academic problem design
English writing

Education

PhD in Mathematics
Doctoral candidate
Master's degree (exceptional depth)

Tools

LaTeX
Markdown
Git

Job description

Role Overview

We are seeking expert mathematicians to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core mathematics domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities.

You will be assigned one of two task types:

  • Question Authoring — Create original, challenging multiple-choice questions in your area of mathematical expertise, rate their difficulty, and submit them for review.
  • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made.
Mathematics Domains Covered

Signal Processing, Financial Mathematics & Actuarial Science, Mathematical Economics, Mathematical Modeling of Ecological & Biological Systems, Mathematical Programming & Combinatorial Optimization, Geomathematics & Climate Modeling.

Key Responsibilities
  • Author original math questions that test deep conceptual understanding, not surface-level recall
  • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement
  • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above)
  • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers
  • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format
  • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories)
  • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made
Ideal Qualifications
  • PhD or doctoral candidate in Mathematics, Applied Mathematics, Statistics, or a closely related field
  • Master's degree considered for candidates with exceptional depth in a specific subdomain
  • Strong command of graduate-level mathematical concepts and formal proof writing
  • Experience with rigorous academic problem design or mathematical competition writing is a strong plus
  • Excellent written English and ability to express complex ideas clearly and concisely
More About the Opportunity
  • Expected commitment: 10+ hours/week
  • Asynchronous, fully remote work
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Mathematics PhD — AI Assessment Question Architect
Remote Mathematics PhD — AI Assessment Question Architect

Obsidian • New York (NY)

On-site
USD 90,000 - 130,000
Remote Mathematics PhD - AI Assessment Content Expert
Remote Mathematics PhD - AI Assessment Content Expert

Mercor • New York (NY)

On-site
USD 83,000 - 179,000
Remote work
Flexible hours
Collaborative AI research environment
Engineering PhD - Benchmark Specialist
Engineering PhD - Benchmark Specialist

Mercor • New York (NY)

On-site
USD 55,000 - 110,000
Remote Mathematics Researcher for AI Assessment Content
Remote Mathematics Researcher for AI Assessment Content

Obsidian • New York (NY)

Remote
USD 120,000 - 180,000
Remote Math Benchmark Architect for AI Research
Remote Math Benchmark Architect for AI Research

Mercor • United States

Remote
USD 90,000 - 150,000
Economics PhD - Benchmark Specialist
Economics PhD - Benchmark Specialist

Mercor • New York (NY)

On-site
USD 83,000 - 138,000
Economics PhD - Benchmark Specialist - AI Trainer
Economics PhD - Benchmark Specialist - AI Trainer

Mercor • Philadelphia

On-site
USD 83,000 - 152,000
Remote work
Flexible schedule
Mathematics PhD (Graduate or Candidate)
Mathematics PhD (Graduate or Candidate)

Weekday AI (YC W21) • United States

Remote
USD 83,000 - 165,000
Mathematics Specialist
Mathematics Specialist

SME Careers • Idaho

On-site
USD 80,000 - 120,000
AI Assessment Specialist - PhD
AI Assessment Specialist - PhD

Mercor • New York (NY)

On-site
USD 55,000 - 124,000
Fully remote