AI Benchmark CS Question Author (Remote Contractor)

Weekday 1

United States

Remote

USD 91,000 - 116,000

Part time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Fully remote

Job summary

Weekday 1 is seeking experienced computer science professionals to author and review high-quality academic assessment content for an AI research initiative. You will develop and validate rigorous multiple-choice questions across CS domains, assess solution quality, and help establish gold-standard benchmarks for evaluating advanced AI systems.

You will contribute through two primary task types: Question Authoring and Question Verification, ensuring accuracy, clarity, and rigor.

Qualifications

  • PhD or doctoral candidacy in CS, EE, or related field.
  • Strong command of graduate-level CS theory, algorithms, systems, software engineering, or ML.
  • Demonstrated depth in at least one listed CS domain.
  • Excellent written English and ability to communicate complex concepts clearly.

Responsibilities

  • Create original CS questions evaluating deep conceptual understanding and problem-solving.
  • Ensure questions are unambiguous, self-contained, and technically accurate.
  • Classify questions by difficulty: Medium, Hard, Expert.
  • Provide one correct answer with nine plausible distractors.
  • Develop clear solution explanations showing reasoning and principles.
  • Provide 1–5 references per question from reputable sources.
  • Identify and explain edits for verification tasks.
  • Apply consistent standards for benchmark quality and reproducibility.

Skills

Academic writing
Technical writing
Question authoring
Question verification

Education

PhD in Computer Science
Master’s in Computer Science

Job description

Weekday 1 is seeking experienced computer science professionals to author and review high-quality academic assessment content for an AI research initiative. You will develop and validate rigorous multiple-choice questions across CS domains, assess solution quality, and help establish gold-standard benchmarks for evaluating advanced AI systems.

You will contribute through two primary task types: Question Authoring and Question Verification, ensuring accuracy, clarity, and rigor.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AI Benchmark Engineer — PhD-Level Question Author
Remote AI Benchmark Engineer — PhD-Level Question Author

Obsidian • Detroit (MI)

On-site
USD 90,000 - 130,000
AI Math Benchmark Architect — Remote
AI Math Benchmark Architect — Remote

Weekday 1 • United States

Remote
USD 84,000 - 106,000
AI Assessment Specialist - PhD
AI Assessment Specialist - PhD

Mercor • New York (NY)

On-site
USD 55,000 - 124,000
Fully remote
AI Assessment Specialist - PhD - AI Trainer
AI Assessment Specialist - PhD - AI Trainer

Mercor • Philadelphia

On-site
USD 50,000 - 70,000
Remote AI Assessment Architect
Remote AI Assessment Architect

Mercor • New York (NY)

On-site
USD 55,000 - 124,000
Fully remote
AI Assessment Specialist — PhD Trainer (Remote)
AI Assessment Specialist — PhD Trainer (Remote)

Mercor • Philadelphia

On-site
USD 50,000 - 70,000
Remote AI Engineering Assessment Author (PhD)
Remote AI Engineering Assessment Author (PhD)

Obsidian • San Francisco (CA)

On-site
USD 83,000 - 124,000
Applied Computer Science Benchmark Specialist
Applied Computer Science Benchmark Specialist

Weekday 1 • United States

Remote
USD 91,000 - 116,000
Fully remote
Remote AI Benchmark Engineer (PhD)
Remote AI Benchmark Engineer (PhD)

Obsidian • New York (NY)

On-site
USD 90,000 - 130,000
AI Benchmark Engineer — PhD (Remote & Flexible)
AI Benchmark Engineer — PhD (Remote & Flexible)

Mercor • New York (NY)

On-site
USD 55,000 - 110,000