AI Benchmark Scientist — Computational Mathematics

Mercor

Greater London

Hybrid

GBP 55,000 - 96,000

Part time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Mercor is seeking PhD and Master’s scientists to author AI evaluation tasks (Sci Code) for a 6-week, part-time engagement in London. You will design original, executable research problems for frontier models and calibrate scoring against model failures.

Responsibilities include sourcing materials, writing prompts, building grading criteria, and testing runs in a pull-request workflow with Docker/Git. Start date is immediate; ~20+ hours per week with flexible scheduling and remote-friendly setup

Qualifications

  • PhD in mathematics, applied mathematics, computational mathematics, or closely related field.
  • Depth in at least two of: numerical linear algebra, computational mechanics, computational finance.
  • Working proficiency in Python or R for scientific computing.
  • Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design.
  • Write scientific prompts based on the input
  • Build the grading criteria that define a correct answer
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed

Skills

Python
R

Education

PhD in Mathematics
Master's in Mathematics or related field

Tools

Git/GitHub
Docker

Job description

Mercor is seeking PhD and Master’s scientists to author AI evaluation tasks (Sci Code) for a 6-week, part-time engagement in London. You will design original, executable research problems for frontier models and calibrate scoring against model failures.

Responsibilities include sourcing materials, writing prompts, building grading criteria, and testing runs in a pull-request workflow with Docker/Git. Start date is immediate; ~20+ hours per week with flexible scheduling and remote-friendly setup

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Benchmark Scientist: Biochemistry & Genetics
AI Benchmark Scientist: Biochemistry & Genetics

Obsidian • Greater London

Remote
GBP 40,000 - 50,000
AI Benchmark Scientist: Biology PhD Coder
AI Benchmark Scientist: Biology PhD Coder

Mercor • Greater London

On-site
GBP 55,000 - 83,000
AI Evaluation Scientist, Math PhD — Frontier Benchmark
AI Evaluation Scientist, Math PhD — Frontier Benchmark

Mercor • Greater London

On-site
GBP 69,000 - 124,000
Physics PhD — AI Benchmark Scientist
Physics PhD — AI Benchmark Scientist

Mercor • Greater London

On-site
GBP 39,000 - 67,000
Physics AI Benchmark Scientist: Prompt Design & Evaluation
Physics AI Benchmark Scientist: Prompt Design & Evaluation

Mercor • Greater London

Remote
GBP 11,000 - 18,000
AI Benchmark Scientist - Quantum & Computational Chemistry
AI Benchmark Scientist - Quantum & Computational Chemistry

Obsidian • Greater London

Remote
GBP 41,000 - 69,000
Computational Mathematician: AI Benchmark Problem Designer
Computational Mathematician: AI Benchmark Problem Designer

Mercor • Greater London

Remote
GBP 109,000 - 218,000
Computational Chemist: AI Benchmark Developer (6-Week)
Computational Chemist: AI Benchmark Developer (6-Week)

Mercor • Greater London

Remote
GBP 6,612,000 - 9,919,000
Quantum Chemistry AI Benchmark Architect (Part-Time, 6 Wks)
Quantum Chemistry AI Benchmark Architect (Part-Time, 6 Wks)

Mercor • Greater London

On-site
GBP 41,000 - 69,000
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Mercor • Greater London

On-site
GBP 69,000 - 124,000