Remote AI Benchmark Engineer — PhD-Level Question Author

Obsidian

Detroit (MI)

On-site

USD 90,000 - 130,000

Part time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Obsidian is seeking expert engineers to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core engineering domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities.

You will be assigned one of two task types: Question Authoring or Question Verification, both remote and asynchronous, with a commitment of 10+ hours per week.

Qualifications

  • PhD or doctoral candidate in Engineering or closely related field.
  • Excellent written English and ability to express complex ideas clearly.
  • Experience with academic assessment content is a plus.

Responsibilities

  • Author original engineering questions that test deep conceptual understanding, not surface-level recall.
  • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement.
  • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above).
  • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers.
  • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format.
  • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories).
  • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made.

Skills

Graduate-level engineering principles

Education

PhD
Master's degree (exceptional depth)

Job description

Obsidian is seeking expert engineers to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core engineering domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities.

You will be assigned one of two task types: Question Authoring or Question Verification, both remote and asynchronous, with a commitment of 10+ hours per week.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI Benchmark Engineer (PhD)
Remote AI Benchmark Engineer (PhD)

Obsidian • New York (NY)

On-site
USD 90,000 - 130,000
Remote AI Assessment Architect (PhD) & Trainer
Remote AI Assessment Architect (PhD) & Trainer

Obsidian • Philadelphia

On-site
USD 69,000 - 96,000
Remote Mathematics PhD — AI Assessment Question Architect
Remote Mathematics PhD — AI Assessment Question Architect

Obsidian • New York (NY)

On-site
USD 90,000 - 130,000
AI Benchmark Engineer — PhD (Remote & Flexible)
AI Benchmark Engineer — PhD (Remote & Flexible)

Mercor • New York (NY)

On-site
USD 55,000 - 110,000
Remote AI Psychology Assessment Architect
Remote AI Psychology Assessment Architect

Obsidian • New York (NY)

On-site
USD 55,000 - 110,000
Remote: Engineering Benchmark & Question Authoring
Remote: Engineering Benchmark & Question Authoring

Mercor • United States

Remote
USD 70,000 - 110,000
Fully remote
Remote Law Content Author for AI Assessment & Benchmarks
Remote Law Content Author for AI Assessment & Benchmarks

Obsidian • New York (NY)

Remote
USD 60,000 - 90,000
Fully remote
Asynchronous work
10+ hours/week
Remote AI Philosophy Benchmark Architect
Remote AI Philosophy Benchmark Architect

Mercor • United States

Remote
USD 60,000 - 85,000
Remote Psychology PhD Assessment Specialist - AI Content & QA
Remote Psychology PhD Assessment Specialist - AI Content & QA

Obsidian • New York (NY)

On-site
USD 83,000 - 131,000
Remote AI Assessment Architect
Remote AI Assessment Architect

Mercor • New York (NY)

On-site
USD 55,000 - 124,000
Fully remote