Applied Computer Science Benchmark Specialist

Weekday 1

United States

Remote

USD 91,000 - 116,000

Part time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Fully remote

Job summary

Weekday 1 is seeking experienced computer science professionals to author and review high-quality academic assessment content for an AI research initiative. You will develop and validate rigorous multiple-choice questions across CS domains, assess solution quality, and help establish gold-standard benchmarks for evaluating advanced AI systems.

You will contribute through two primary task types: Question Authoring and Question Verification, ensuring accuracy, clarity, and rigor.

Qualifications

  • PhD or doctoral candidacy in CS, EE, or related field.
  • Strong command of graduate-level CS theory, algorithms, systems, software engineering, or ML.
  • Demonstrated depth in at least one listed CS domain.
  • Excellent written English and ability to communicate complex concepts clearly.

Responsibilities

  • Create original CS questions evaluating deep conceptual understanding and problem-solving.
  • Ensure questions are unambiguous, self-contained, and technically accurate.
  • Classify questions by difficulty: Medium, Hard, Expert.
  • Provide one correct answer with nine plausible distractors.
  • Develop clear solution explanations showing reasoning and principles.
  • Provide 1–5 references per question from reputable sources.
  • Identify and explain edits for verification tasks.
  • Apply consistent standards for benchmark quality and reproducibility.

Skills

Academic writing
Technical writing
Question authoring
Question verification

Education

PhD in Computer Science
Master’s in Computer Science

Job description

This role is for one of our clients

Compensation: $66 - $84 per hour

We are seeking experienced computer science professionals to author and review high-quality academic assessment content for an AI research initiative. In this role, you will develop and validate rigorous multiple-choice questions across a broad range of computer science domains, assess solution quality, and help establish gold-standard benchmarks for evaluating advanced AI systems.

You will contribute through one of two primary task types:

Question Authoring - Develop original, challenging multiple-choice questions within your area of computer science expertise, assess their difficulty, and submit them for review.

Question Verification - Review existing questions for technical accuracy, clarity, completeness, and rigor. Make necessary edits, assess difficulty, and document the rationale behind your changes.

Requirements

Computer Science Domains
  • Accelerator / GPU Kernel Engineering
  • Formal Methods & Automated Reasoning
  • Computer Architecture & Accelerators
  • Distributed Systems
  • DevOps & Site Reliability Engineering
  • Data Engineering & Databases
  • Cloud Computing & Infrastructure
  • Operating Systems & Systems Kernel
  • Machine Learning Engineering
  • Web & API Development
  • Embedded Systems Engineering
  • Computer Graphics & Game Development
  • Mobile Engineering
Key Responsibilities
  • Create original computer science questions that evaluate deep conceptual understanding, technical reasoning, and problem-solving rather than surface-level recall.
  • Ensure every question is unambiguous, self-contained, technically accurate, and sufficiently specified for a qualified expert to solve.
  • Classify questions by difficulty:
    • Medium: Introductory undergraduate level
    • Hard: Advanced undergraduate level
    • Expert: Postgraduate level and above
  • Provide one correct answer alongside nine plausible but subtly incorrect alternatives designed to distinguish strong technical reasoning from superficial knowledge.
  • Develop clear, structured solution explanations that demonstrate the reasoning and technical principles required to reach the correct answer.
  • Provide 1-5 authoritative references per question, drawing from peer-reviewed research, academic publications, university resources, and other reputable technical sources.
  • For verification assignments, identify issues related to correctness, clarity, completeness, precision, or solvability and clearly explain the reasoning behind any recommended edits.
  • Apply consistent standards when evaluating questions and solutions to ensure benchmark quality and reproducibility.
Ideal Qualifications
  • PhD or doctoral candidacy in Computer Science, Electrical Engineering, Computer Engineering, or a closely related discipline.
  • A Master's degree may be considered for candidates with exceptional expertise in a specialized computer science domain.
  • Strong command of graduate-level computer science theory, algorithms, systems, software engineering, architecture, and/or machine learning.
  • Demonstrated depth in one or more of the listed technical domains.
  • Research publications, substantial industry experience at leading technology organizations, systems engineering experience, or competitive programming experience is a strong plus.
  • Excellent written English and the ability to communicate complex technical concepts clearly, accurately, and concisely.
  • Strong attention to detail and the ability to distinguish technically valid solutions from plausible but incorrect approaches.
More About the Opportunity
  • Expected commitment: 10+ hours per week
  • Fully remote and asynchronous
  • Flexible scheduling based on project requirements
  • Opportunity to contribute to the development of high-quality benchmarks for evaluating advanced AI systems
  • Strong contributors may be considered for additional review, evaluation, or subject-matter expert opportunities
Equal Opportunity

We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations throughout the application and engagement process.

Contract and Payment Terms
  • Engagement will be on an independent contractor basis.
  • This is a fully remote opportunity that can be completed on your own schedule.
  • Projects may be extended, shortened, or concluded early depending on project requirements and performance.
  • Work will not require access to confidential or proprietary information belonging to any current or former employer, client, or institution.
  • Payments are made weekly through Stripe or Wise, based on services rendered.
  • H-1B and STEM OPT candidates are not eligible for this opportunity at this time.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Applied Mathematics Benchmark Specialist
Applied Mathematics Benchmark Specialist

Weekday 1 • United States

Remote
USD 84,000 - 106,000
Applied Legal Benchmark Specialist
Applied Legal Benchmark Specialist

Weekday 1 • United States

Remote
USD 114,000 - 145,000
AI Assessment Specialist - PhD
AI Assessment Specialist - PhD

Mercor • New York (NY)

On-site
USD 55,000 - 124,000
Fully remote
Software Engineering Expert
Software Engineering Expert

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Fully remote
Weekly payments
AI Assessment Specialist - PhD - AI Trainer
AI Assessment Specialist - PhD - AI Trainer

Mercor • Philadelphia

On-site
USD 50,000 - 70,000
AI Benchmark CS Question Author (Remote Contractor)
AI Benchmark CS Question Author (Remote Contractor)

Weekday 1 • United States

Remote
USD 91,000 - 116,000
Fully remote
Applied History & Political Science Benchmark Specialist
Applied History & Political Science Benchmark Specialist

Weekday 1 • United States

Remote
USD 61,000 - 77,000
Data Science & Quantitative Analysis Expert
Data Science & Quantitative Analysis Expert

Weekday 1 • United States

Remote
USD 109,000 - 164,000
Fully remote
Weekly payments
STEM Researcher - Computational Fields
STEM Researcher - Computational Fields

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Engineering PhD - Benchmark Specialist
Engineering PhD - Benchmark Specialist

Mercor • New York (NY)

On-site
USD 55,000 - 110,000