AI Benchmark Scientist - Mathematics PhD (6-Week Project)

Weekday 1

United States

On-site

USD 83,000 - 110,000

Part time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Weekday 1 is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark for scientific computing. You will author original, executable research problems that today's frontier models cannot solve.

You will source material, write scientific prompts, and build grading criteria. The role requires calibration against frontier models, with a 6-week, part-time commitment of 20+ hours per week and immediate start.

Qualifications

  • PhD in mathematics, applied mathematics, computational mathematics, or a closely related field.
  • Demonstrated depth in at least two of: numerical linear algebra, computational mechanics, computational finance.
  • Working proficiency in Python or R for scientific computing.
  • Comfortable with Git/GitHub and running code in Docker - authoring runs through a pull-request workflow with automated quality checks.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design.
  • Write scientific prompts based on the input
  • Build the grading criteria that define a correct answer
  • Calibrate against frontier models - a task ships only when strong models fail it more often than they succeed

Skills

Python
R
Git/GitHub
Docker

Education

PhD in mathematics, applied mathematics, or computational mathematics

Tools

Docker
GitHub

Job description

Weekday 1 is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark for scientific computing. You will author original, executable research problems that today's frontier models cannot solve.

You will source material, write scientific prompts, and build grading criteria. The role requires calibration against frontier models, with a 6-week, part-time commitment of 20+ hours per week and immediate start.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Benchmark Scientist: Physics PhD in Scientific Computing
AI Benchmark Scientist: Physics PhD in Scientific Computing

Obsidian • Dallas (TX)

On-site
USD 9,919,000 - 14,878,000
AI Benchmark Architect: Computational Mathematician
AI Benchmark Architect: Computational Mathematician

Mercor • San Francisco (CA)

Remote
USD 83,000 - 165,000
Remote AI Benchmark Scientist (Physics PhD)
Remote AI Benchmark Scientist (Physics PhD)

Weekday 1 • United States

Remote
USD 83,000 - 110,000
Bio PhD AI Benchmark Architect (Part-Time, 6-Week Project)
Bio PhD AI Benchmark Architect (Part-Time, 6-Week Project)

Mercor • San Diego (CA)

On-site
USD 69,000 - 103,000
AI Benchmark Scientist - Materials Science & Modeling (PhD)
AI Benchmark Scientist - Materials Science & Modeling (PhD)

Obsidian • New York (NY)

On-site
USD 83,000 - 138,000
Chemistry PhD AI Benchmark Architect (Remote)
Chemistry PhD AI Benchmark Architect (Remote)

Weekday 1 • United States

Remote
USD 83,000 - 110,000
Biology PhD Scientist for AI Benchmark Coding (6 Weeks, Part-Time)
Biology PhD Scientist for AI Benchmark Coding (6 Weeks, Part-Time)

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 138,000
6-Week Part-Time Materials Scientist for AI Benchmarks
6-Week Part-Time Materials Scientist for AI Benchmarks

Mercor • San Francisco (CA)

Remote
USD 150,000 - 190,000
AI Benchmark Scientist for Frontier Physics Computing
AI Benchmark Scientist for Frontier Physics Computing

Mercor • Dallas (TX)

On-site
USD 83,000 - 124,000
AI Benchmark Architect for Scientific Computing
AI Benchmark Architect for Scientific Computing

Obsidian • San Francisco (CA)

Remote
USD 96,000 - 138,000
6-week engagement
Part-time 20+ hrs/week
Immediate start