Remote AI Benchmark Engineer (PhD)

Obsidian

New York (NY)

On-site

USD 90,000 - 130,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Obsidian is seeking expert engineers to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core engineering domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities.

You will be assigned one of two task types: Question Authoring — Create original, challenging multiple-choice questions in your area of engineering expertise, rate their

Qualifications

  • PhD or doctoral candidate in Engineering or a closely related field.
  • Master's degree considered for candidates with exceptional depth in a specific subdomain.
  • Excellent written English and ability to express complex ideas clearly.
  • Professional engineering licensure (PE) or industry experience is a strong plus.
  • Strong command of graduate-level engineering principles, applied mathematics, and domain-specific standards.

Responsibilities

  • Author original engineering questions that test deep conceptual understanding, not surface-level recall
  • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement
  • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above)
  • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers
  • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format
  • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories)
  • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made

Skills

Question authoring
Question verification
Academic writing
Peer review

Education

PhD in Engineering
Doctoral candidate in Engineering
Master's degree in a relevant field

Job description

Obsidian is seeking expert engineers to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core engineering domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities.

You will be assigned one of two task types: Question Authoring — Create original, challenging multiple-choice questions in your area of engineering expertise, rate their

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI Benchmark Engineer — PhD-Level Question Author
Remote AI Benchmark Engineer — PhD-Level Question Author

Obsidian • Detroit (MI)

On-site
USD 90,000 - 130,000
Remote AI Assessment Architect (PhD) & Trainer
Remote AI Assessment Architect (PhD) & Trainer

Obsidian • Philadelphia

On-site
USD 69,000 - 96,000
AI Benchmark Engineer — PhD (Remote & Flexible)
AI Benchmark Engineer — PhD (Remote & Flexible)

Mercor • New York (NY)

On-site
USD 55,000 - 110,000
Remote AI Psychology Assessment Architect
Remote AI Psychology Assessment Architect

Obsidian • New York (NY)

On-site
USD 55,000 - 110,000
Remote: Engineering Benchmark & Question Authoring
Remote: Engineering Benchmark & Question Authoring

Mercor • United States

Remote
USD 70,000 - 110,000
Fully remote
Remote Mathematics PhD — AI Assessment Question Architect
Remote Mathematics PhD — AI Assessment Question Architect

Obsidian • New York (NY)

On-site
USD 90,000 - 130,000
Remote AI Business & Commerce Assessment Architect
Remote AI Business & Commerce Assessment Architect

Obsidian • New York (NY)

Remote
USD 65,000 - 95,000
Remote Law Content Author for AI Assessment & Benchmarks
Remote Law Content Author for AI Assessment & Benchmarks

Obsidian • New York (NY)

Remote
USD 60,000 - 90,000
Fully remote
Asynchronous work
10+ hours/week
Remote Mathematics Researcher for AI Assessment Content
Remote Mathematics Researcher for AI Assessment Content

Obsidian • New York (NY)

Remote
USD 120,000 - 180,000
Remote Psychology PhD Assessment Specialist - AI Content & QA
Remote Psychology PhD Assessment Specialist - AI Content & QA

Obsidian • New York (NY)

On-site
USD 83,000 - 131,000