AI Evaluation Scientist, Math PhD — Frontier Benchmark

Mercor

Greater London

On-site

GBP 69,000 - 124,000

Part time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Mercor is hiring PhD and Master’s scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today’s frontier models cannot solve.

Engage with leading AI labs to build benchmarks for scientific computing and craft robust grading criteria. You will work on multiple subdomains with a coding focus, using Python or R, and Docker-based workflows. Start date is immediate for a 6-week, part-time engagement.

Qualifications

  • PhD in mathematics, applied mathematics, computational mathematics, or a closely related field.
  • Depth in at least two of the following: numerical linear algebra, computational mechanics, computational finance.
  • Working proficiency in Python or R for scientific computing.
  • Comfortable with Git/GitHub and Docker; authoring runs through a pull-request workflow with automated quality checks.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design
  • Write scientific prompts based on the input
  • Build the grading criteria that define a correct answer
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed

Skills

Python
R
Git/GitHub
Docker
Numerical linear algebra
Computational mechanics
Computational finance

Education

PhD in Mathematics

Job description

Mercor is hiring PhD and Master’s scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today’s frontier models cannot solve.

Engage with leading AI labs to build benchmarks for scientific computing and craft robust grading criteria. You will work on multiple subdomains with a coding focus, using Python or R, and Docker-based workflows. Start date is immediate for a 6-week, part-time engagement.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Mercor • Greater London

On-site
GBP 69,000 - 124,000
Physics PhD — AI Benchmark Scientist
Physics PhD — AI Benchmark Scientist

Mercor • Greater London

On-site
GBP 39,000 - 67,000
AI Benchmark Scientist: Biology PhD Coder
AI Benchmark Scientist: Biology PhD Coder

Mercor • Greater London

On-site
GBP 55,000 - 83,000
Computational Mathematician: AI Benchmark Problem Designer
Computational Mathematician: AI Benchmark Problem Designer

Mercor • Greater London

Remote
GBP 109,000 - 218,000
AI Benchmark Scientist: Biochemistry & Genetics
AI Benchmark Scientist: Biochemistry & Genetics

Obsidian • Greater London

Remote
GBP 40,000 - 50,000
Quantum Computing AI Benchmark Designer
Quantum Computing AI Benchmark Designer

Mercor • Greater London

On-site
GBP 273,000 - 455,000
Physics AI Benchmark Scientist: Prompt Design & Evaluation
Physics AI Benchmark Scientist: Prompt Design & Evaluation

Mercor • Greater London

Remote
GBP 11,000 - 18,000
Biology PhD AI Benchmark Coder
Biology PhD AI Benchmark Coder

Obsidian • Greater London

On-site
GBP 4,959,000 - 6,612,000
Quantum Chemistry AI Research Engineer (Part-Time)
Quantum Chemistry AI Research Engineer (Part-Time)

Obsidian • Greater London

On-site
GBP 55,000 - 83,000
Physics PhD - Scientific Computing Expert
Physics PhD - Scientific Computing Expert

Mercor • Greater London

On-site
GBP 39,000 - 67,000