AI Benchmark Engineer — PhD (Remote & Flexible)

Mercor

New York (NY)

On-site

USD 55,000 - 110,000

Part time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Mercor is seeking expert engineers to author and review high-quality academic assessment content for an AI research initiative. You will craft and verify rigorous MCQs across core engineering domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities.

You will work asynchronously from anywhere, committing roughly 10+ hours per week. The role emphasizes deep conceptual problem design, precise wording, and clear explanations suitable for

Qualifications

  • PhD or doctoral candidate in Engineering required or highly preferred.
  • Excellent written English and concise technical communication.
  • Strong command of graduate-level engineering principles and applied mathematics.

Responsibilities

  • Author original engineering questions that test deep conceptual understanding.
  • Rate difficulty and provide 1 correct answer plus 9 plausible distractors.
  • Review and edit pre-written questions for accuracy and rigor.
  • Provide step-by-step Chain-of-Thought solutions in markdown format.
  • Supply 1–5 academic references per question.
  • Flag clarity and solvability issues when verifying questions.

Skills

Graduate-level engineering principles
Technical writing
English proficiency

Education

PhD or doctoral candidate in Engineering
Master's degree in engineering

Job description

Mercor is seeking expert engineers to author and review high-quality academic assessment content for an AI research initiative. You will craft and verify rigorous MCQs across core engineering domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities.

You will work asynchronously from anywhere, committing roughly 10+ hours per week. The role emphasizes deep conceptual problem design, precise wording, and clear explanations suitable for

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Engineering PhD — AI Benchmark Question Architect
Engineering PhD — AI Benchmark Question Architect

Mercor • San Francisco (CA)

On-site
USD 83,000 - 165,000
Fully remote
Asynchronous work
Remote Engineering PhD Assessment Specialist
Remote Engineering PhD Assessment Specialist

Mercor • New York (NY)

On-site
USD 48,000 - 96,000
Remote AI Benchmark Engineer — PhD-Level Question Author
Remote AI Benchmark Engineer — PhD-Level Question Author

Obsidian • Detroit (MI)

On-site
USD 90,000 - 130,000
Remote AI Benchmark Engineer (PhD)
Remote AI Benchmark Engineer (PhD)

Obsidian • New York (NY)

On-site
USD 90,000 - 130,000
AI CS Benchmark Engineer (Remote)
AI CS Benchmark Engineer (Remote)

1000scholars • United States

Remote
USD 90,000 - 140,000
AI Assessment Specialist — PhD Trainer (Remote)
AI Assessment Specialist — PhD Trainer (Remote)

Mercor • Philadelphia

On-site
USD 50,000 - 70,000
AI Math Benchmark Architect — Remote
AI Math Benchmark Architect — Remote

Weekday 1 • United States

Remote
USD 84,000 - 106,000
Economics PhD - Benchmark & Assessment Specialist (Remote)
Economics PhD - Benchmark & Assessment Specialist (Remote)

Mercor • Philadelphia

On-site
USD 83,000 - 152,000
Remote work
Flexible schedule
Remote AI Engineering Assessment Author (PhD)
Remote AI Engineering Assessment Author (PhD)

Obsidian • San Francisco (CA)

On-site
USD 83,000 - 124,000
Remote Applied Physics Benchmark Specialist for AI Research
Remote Applied Physics Benchmark Specialist for AI Research

1000scholars • United States

Remote
USD 90,000 - 130,000