Applied Computer Science Benchmark Specialist

Weekday AI

United States

Remote

USD 91,000 - 116,000

Part time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Weekday AI is seeking experienced computer science professionals to author and review high-quality academic assessment content for an AI research initiative. You will develop and validate rigorous multiple-choice questions across CS domains, assess solution quality, and help establish gold-standard benchmarks for evaluating advanced AI systems.

You will contribute as a Question Authoring or Question Verification role, developing original questions or reviewing existing ones for accuracy,

Qualifications

  • PhD or doctoral candidacy in Computer Science, Electrical Engineering, Computer Engineering, or closely related discipline.
  • Strong command of graduate-level CS theory, algorithms, systems, software engineering, architecture, and/or machine learning.
  • Demonstrated depth in one or more of the listed technical domains.

Responsibilities

  • Create original CS questions that evaluate deep conceptual understanding and problem solving.
  • Verify existing questions for technical accuracy, clarity, completeness, and rigor; edit as needed and document reasoning.
  • Provide clear, structured solution explanations and references for each question.
  • Identify issues in verification assignments and explain recommended edits clearly.
  • Apply consistent standards to ensure benchmark quality and reproducibility.

Skills

PhD or doctoral candidacy
Written English proficiency
Detail-oriented
Clear technical communication
Graduate-level CS theory

Education

PhD in Computer Science or related
Master's degree in CS (considered)

Job description

This role is for one of our clients

Compensation: $66 - $84 per hour

We are seeking experienced computer science professionals to author and review high-quality academic assessment content for an AI research initiative. In this role, you will develop and validate rigorous multiple-choice questions across a broad range of computer science domains, assess solution quality, and help establish gold-standard benchmarks for evaluating advanced AI systems.

You will contribute through one of two primary task types:

Question Authoring — Develop original, challenging multiple-choice questions within your area of computer science expertise, assess their difficulty, and submit them for review.

Question Verification — Review existing questions for technical accuracy, clarity, completeness, and rigor. Make necessary edits, assess difficulty, and document the rationale behind your changes.

Computer Science Domains
  • Accelerator / GPU Kernel Engineering
  • Formal Methods & Automated Reasoning
  • Computer Architecture & Accelerators
  • Distributed Systems
  • DevOps & Site Reliability Engineering
  • Data Engineering & Databases
  • Cloud Computing & Infrastructure
  • Operating Systems & Systems Kernel
  • Machine Learning Engineering
  • Web & API Development
  • Embedded Systems Engineering
  • Computer Graphics & Game Development
  • Mobile Engineering
Key Responsibilities
  • Create original computer science questions that evaluate deep conceptual understanding, technical reasoning, and problem-solving rather than surface-level recall.
  • Ensure every question is unambiguous, self-contained, technically accurate, and sufficiently specified for a qualified expert to solve.
  • Classify questions by difficulty:
    • Medium: Introductory undergraduate level
    • Hard: Advanced undergraduate level
    • Expert: Postgraduate level and above
  • Provide one correct answer alongside nine plausible but subtly incorrect alternatives designed to distinguish strong technical reasoning from superficial knowledge.
  • Develop clear, structured solution explanations that demonstrate the reasoning and technical principles required to reach the correct answer.
  • Provide 1–5 authoritative references per question, drawing from peer-reviewed research, academic publications, university resources, and other reputable technical sources.
  • For verification assignments, identify issues related to correctness, clarity, completeness, precision, or solvability and clearly explain the reasoning behind any recommended edits.
  • Apply consistent standards when evaluating questions and solutions to ensure benchmark quality and reproducibility.
Ideal Qualifications
  • PhD or doctoral candidacy in Computer Science, Electrical Engineering, Computer Engineering, or a closely related discipline.
  • A Master's degree may be considered for candidates with exceptional expertise in a specialized computer science domain.
  • Strong command of graduate-level computer science theory, algorithms, systems, software engineering, architecture, and/or machine learning.
  • Demonstrated depth in one or more of the listed technical domains.
  • Research publications, substantial industry experience at leading technology organizations, systems engineering experience, or competitive programming experience is a strong plus.
  • Excellent written English and the ability to communicate complex technical concepts clearly, accurately, and concisely.
  • Strong attention to detail and the ability to distinguish technically valid solutions from plausible but incorrect approaches.
More About the Opportunity
  • Expected commitment: 10+ hours per week
  • Fully remote and asynchronous
  • Flexible scheduling based on project requirements
  • Opportunity to contribute to the development of high-quality benchmarks for evaluating advanced AI systems
  • Strong contributors may be considered for additional review, evaluation, or subject-matter expert opportunities
Equal Opportunity

We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations throughout the application and engagement process.

Contract and Payment Terms
  • Engagement will be on an independent contractor basis.
  • This is a fully remote opportunity that can be completed on your own schedule.
  • Projects may be extended, shortened, or concluded early depending on project requirements and performance.
  • Work will not require access to confidential or proprietary information belonging to any current or former employer, client, or institution.
  • Payments are made weekly through Stripe or Wise, based on services rendered.
  • H-1B and STEM OPT candidates are not eligible for this opportunity at this time.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Applied Computer Science Benchmark Specialist
Applied Computer Science Benchmark Specialist

Weekday 1 • United States

Remote
USD 91,000 - 116,000
Fully remote
Applied Health & Medicine Benchmark Specialist
Applied Health & Medicine Benchmark Specialist

Weekday 1 • United States

Remote
USD 129,000 - 164,000
Fully Remote
Independent Contractor
Weekly payments
Applied Health & Medicine Benchmark Specialist
Applied Health & Medicine Benchmark Specialist

Weekday AI • United States

Remote
USD 129,000 - 164,000
Fully Remote
Independent Contractor
Applied Legal Benchmark Specialist
Applied Legal Benchmark Specialist

Weekday AI • United States

Remote
USD 114,000 - 145,000
Applied Mathematics Benchmark Specialist
Applied Mathematics Benchmark Specialist

Weekday 1 • United States

Remote
USD 84,000 - 106,000
AI Assessment Specialist - PhD
AI Assessment Specialist - PhD

Mercor • New York (NY)

On-site
USD 55,000 - 124,000
Fully remote
AI Assessment Specialist - PhD - AI Trainer
AI Assessment Specialist - PhD - AI Trainer

Mercor • Philadelphia

On-site
USD 50,000 - 70,000
Software Engineering Expert
Software Engineering Expert

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Fully remote
Weekly payments
Applied Mathematics Benchmark Specialist
Applied Mathematics Benchmark Specialist

Weekday AI • United States

Remote
USD 84,000 - 106,000
Fully remote
Software Engineering Expert
Software Engineering Expert

Weekday AI • United States

Remote
USD 83,000 - 124,000