Biochemist for AI Benchmark Tasks & Prompt Design

Mercor

San Francisco (CA)

Remote

USD 96,000 - 138,000

Part time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Mercor is hiring PhD and Master’s scientists to author AI evaluation tasks for Sci Code. You will design original, executable research problems that today’s frontier models cannot solve.

Source your own material, write scientific prompts, build grading criteria, and calibrate against frontier models. Tasks ship only when strong models fail it more often than they succeed, with Docker runs and PR-based quality checks.

Qualifications

  • PhD in biology, biological sciences, biochemistry, genetics, ecology, or a closely related field.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design.
  • Write scientific prompts based on the input
  • Build the grading criteria that define a correct answer
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed

Skills

Python
R
Git/GitHub

Education

PhD in Biology/Related Field
Master's in Biology/Related Field

Tools

Docker

Job description

Mercor is hiring PhD and Master’s scientists to author AI evaluation tasks for Sci Code. You will design original, executable research problems that today’s frontier models cannot solve.

Source your own material, write scientific prompts, build grading criteria, and calibrate against frontier models. Tasks ship only when strong models fail it more often than they succeed, with Docker runs and PR-based quality checks.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Benchmark Scientist for Biochemistry & Genomics
AI Benchmark Scientist for Biochemistry & Genomics

Obsidian • San Francisco (CA)

Remote
USD 273,000 - 382,000
Biology PhD: AI Benchmark Scientist & Scientific Coder
Biology PhD: AI Benchmark Scientist & Scientific Coder

Obsidian • San Diego (CA)

On-site
USD 55,000 - 103,000
Materials Scientist: AI Benchmarking & Research Prompts
Materials Scientist: AI Benchmarking & Research Prompts

Obsidian • San Francisco (CA)

Remote
USD 8,266,000 - 12,398,000
Quantum Chemistry AI Researcher & Code Architect
Quantum Chemistry AI Researcher & Code Architect

Obsidian • San Diego (CA)

On-site
USD 83,000 - 124,000
AI Benchmark Architect: Computational Mathematician
AI Benchmark Architect: Computational Mathematician

Mercor • San Francisco (CA)

Remote
USD 83,000 - 165,000
6-Week Part-Time Materials Scientist for AI Benchmarks
6-Week Part-Time Materials Scientist for AI Benchmarks

Mercor • San Francisco (CA)

Remote
USD 150,000 - 190,000
Quantum AI Benchmark Scientist
Quantum AI Benchmark Scientist

Obsidian • New York (NY)

Remote
USD 83,000 - 138,000
Quantum Computing Research Engineer for AI Benchmarks
Quantum Computing Research Engineer for AI Benchmarks

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 165,000
Biology PhD — AI Research Benchmark Architect
Biology PhD — AI Research Benchmark Architect

Obsidian • New York (NY)

On-site
USD 83,000 - 138,000
Quantum Chemistry AI Benchmark Scientist
Quantum Chemistry AI Benchmark Scientist

Obsidian • New York (NY)

On-site
USD 83,000 - 138,000