Remote CS Benchmark Architect for AI Evaluation

Weekday AI

United States

Remote

USD 91,000 - 116,000

Part time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Weekday AI is seeking experienced computer science professionals to author and review high-quality academic assessment content for an AI research initiative. You will develop and validate rigorous multiple-choice questions across CS domains, assess solution quality, and help establish gold-standard benchmarks for evaluating advanced AI systems.

You will contribute as a Question Authoring or Question Verification role, developing original questions or reviewing existing ones for accuracy,

Qualifications

  • PhD or doctoral candidacy in Computer Science, Electrical Engineering, Computer Engineering, or closely related discipline.
  • Strong command of graduate-level CS theory, algorithms, systems, software engineering, architecture, and/or machine learning.
  • Demonstrated depth in one or more of the listed technical domains.

Responsibilities

  • Create original CS questions that evaluate deep conceptual understanding and problem solving.
  • Verify existing questions for technical accuracy, clarity, completeness, and rigor; edit as needed and document reasoning.
  • Provide clear, structured solution explanations and references for each question.
  • Identify issues in verification assignments and explain recommended edits clearly.
  • Apply consistent standards to ensure benchmark quality and reproducibility.

Skills

PhD or doctoral candidacy
Written English proficiency
Detail-oriented
Clear technical communication
Graduate-level CS theory

Education

PhD in Computer Science or related
Master's degree in CS (considered)

Job description

Weekday AI is seeking experienced computer science professionals to author and review high-quality academic assessment content for an AI research initiative. You will develop and validate rigorous multiple-choice questions across CS domains, assess solution quality, and help establish gold-standard benchmarks for evaluating advanced AI systems.

You will contribute as a Question Authoring or Question Verification role, developing original questions or reviewing existing ones for accuracy,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Benchmark CS Question Author (Remote Contractor)
AI Benchmark CS Question Author (Remote Contractor)

Weekday 1 • United States

Remote
USD 91,000 - 116,000
Fully remote
AI CS Benchmark Engineer (Remote)
AI CS Benchmark Engineer (Remote)

1000scholars • United States

Remote
USD 90,000 - 140,000
Remote AI Assessment Architect
Remote AI Assessment Architect

Mercor • New York (NY)

On-site
USD 55,000 - 124,000
Fully remote
Applied Math Benchmark Architect — Remote
Applied Math Benchmark Architect — Remote

Weekday AI • United States

Remote
USD 84,000 - 106,000
Fully remote
Remote AI Benchmark Engineer — PhD-Level Question Author
Remote AI Benchmark Engineer — PhD-Level Question Author

Obsidian • Detroit (MI)

On-site
USD 90,000 - 130,000
Remote AI Benchmark Engineer (PhD)
Remote AI Benchmark Engineer (PhD)

Obsidian • New York (NY)

On-site
USD 90,000 - 130,000
Remote AI Legal Benchmark Architect
Remote AI Legal Benchmark Architect

Weekday 1 • United States

Remote
USD 114,000 - 145,000
Remote AI Biology Benchmark Architect
Remote AI Biology Benchmark Architect

1000scholars • United States

Remote
USD 2,700 - 5,500
Fully remote
Weekly payments via Stripe or Wise
Remote History & Political Science Benchmark Specialist
Remote History & Political Science Benchmark Specialist

Weekday AI • United States

Remote
USD 61,000 - 77,000
AI Math Benchmark Architect — Remote
AI Math Benchmark Architect — Remote

Weekday 1 • United States

Remote
USD 84,000 - 106,000