AI Benchmark Scientist: Biochemistry & Genetics

Obsidian

Greater London

Remote

GBP 40,000 - 50,000

Part time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) as part of a collaboration with leading AI labs to establish a new benchmark for scientific computing. You will author original, executable research problems that frontier models cannot solve.

Based in Greater London, you will source material, write prompts, and build grading criteria, calibrating against frontier models so that tasks ship only when models struggle.

Qualifications

  • PhD in biology, biological sciences, biochemistry, genetics, ecology, or a closely related field.
  • Depth in at least two of ecology, biochemistry, genetics.
  • Working proficiency in Python, R, or another relevant programming language for scientific computing.
  • Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design.
  • Write scientific prompts based on the input
  • Build the grading criteria that define a correct answer
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed

Skills

Python
R
Git/GitHub
Docker

Education

PhD in biology or closely related field

Tools

Python
R

Job description

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) as part of a collaboration with leading AI labs to establish a new benchmark for scientific computing. You will author original, executable research problems that frontier models cannot solve.

Based in Greater London, you will source material, write prompts, and build grading criteria, calibrating against frontier models so that tasks ship only when models struggle.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Benchmark Scientist: Biology PhD Coder
AI Benchmark Scientist: Biology PhD Coder

Mercor • Greater London

On-site
GBP 55,000 - 83,000
AI Benchmark Scientist — Computational Mathematics
AI Benchmark Scientist — Computational Mathematics

Mercor • Greater London

Hybrid
GBP 55,000 - 96,000
AI Benchmark Scientist - Quantum & Computational Chemistry
AI Benchmark Scientist - Quantum & Computational Chemistry

Obsidian • Greater London

Remote
GBP 41,000 - 69,000
Biology PhD AI Benchmark Coder
Biology PhD AI Benchmark Coder

Obsidian • Greater London

On-site
GBP 4,959,000 - 6,612,000
Computational Mathematician: AI Benchmark Problem Designer
Computational Mathematician: AI Benchmark Problem Designer

Mercor • Greater London

Remote
GBP 109,000 - 218,000
Physics PhD — AI Benchmark Scientist
Physics PhD — AI Benchmark Scientist

Mercor • Greater London

On-site
GBP 39,000 - 67,000
Quantum Computing AI Benchmark Designer
Quantum Computing AI Benchmark Designer

Mercor • Greater London

On-site
GBP 273,000 - 455,000
AI Evaluation Scientist, Math PhD — Frontier Benchmark
AI Evaluation Scientist, Math PhD — Frontier Benchmark

Mercor • Greater London

On-site
GBP 69,000 - 124,000
Quantum Chemistry AI Research Engineer (Part-Time)
Quantum Chemistry AI Research Engineer (Part-Time)

Obsidian • Greater London

On-site
GBP 55,000 - 83,000
Physics AI Benchmark Scientist: Prompt Design & Evaluation
Physics AI Benchmark Scientist: Prompt Design & Evaluation

Mercor • Greater London

Remote
GBP 11,000 - 18,000