Bio PhD AI Benchmark Architect (Part-Time, 6-Week Project)

Mercor

San Diego (CA)

On-site

USD 69,000 - 103,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will create original, executable research problems that today’s frontier models cannot solve, sourced from papers, datasets, or your own scenarios.

We expect depth in at least two subdomains like ecology, biochemistry, or genetics, and proficiency in Python or R with Docker/GitHub workflows.

Qualifications

  • PhD in biology, biological sciences, biochemistry, genetics, ecology, or a closely related field.
  • Demonstrated depth in at least two of ecology, biochemistry, genetics.
  • Working proficiency in Python, R, or another relevant programming language for scientific computing.
  • Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design.
  • Write scientific prompts based on the input
  • Build the grading criteria that define a correct answer
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed

Skills

Python
R
Scientific computing

Education

PhD in biology
Masters in biology/related field

Tools

Git/GitHub
Docker

Job description

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will create original, executable research problems that today’s frontier models cannot solve, sourced from papers, datasets, or your own scenarios.

We expect depth in at least two subdomains like ecology, biochemistry, or genetics, and proficiency in Python or R with Docker/GitHub workflows.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Biology PhD Scientist for AI Benchmark Coding (6 Weeks, Part-Time)
Biology PhD Scientist for AI Benchmark Coding (6 Weeks, Part-Time)

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 138,000
Biology PhD — AI Research Benchmark Architect
Biology PhD — AI Research Benchmark Architect

Obsidian • New York (NY)

On-site
USD 83,000 - 138,000
Biology PhD Scientist Coder for AI Benchmark Tasks
Biology PhD Scientist Coder for AI Benchmark Tasks

Mercor • San Francisco (CA)

On-site
USD 83,000 - 124,000
Publications in peer-reviewed journals
Prior scientific software or research‑
AI Evaluation Scientist - Math PhD (6-Week, Part-Time)
AI Evaluation Scientist - Math PhD (6-Week, Part-Time)

Mercor • Miami (FL)

On-site
USD 83,000 - 138,000
Quantum Computing AI Benchmark Scientist (PhD) — Part-Time
Quantum Computing AI Benchmark Scientist (PhD) — Part-Time

Mercor • New York (NY)

On-site
USD 83,000 - 124,000
AI Benchmark Scientist for Biochemistry & Genomics
AI Benchmark Scientist for Biochemistry & Genomics

Obsidian • San Francisco (CA)

Remote
USD 273,000 - 382,000
6-Week Part-Time Materials Scientist for AI Benchmarks
6-Week Part-Time Materials Scientist for AI Benchmarks

Mercor • San Francisco (CA)

Remote
USD 150,000 - 190,000
AI Benchmark Designer: Math PhD & Trainer
AI Benchmark Designer: Math PhD & Trainer

Obsidian • Los Angeles (CA)

On-site
USD 83,000 - 124,000
Biology PhD: AI Benchmark Scientist & Scientific Coder
Biology PhD: AI Benchmark Scientist & Scientific Coder

Obsidian • San Diego (CA)

On-site
USD 55,000 - 103,000
AI Benchmark Architect: Computational Mathematician
AI Benchmark Architect: Computational Mathematician

Mercor • San Francisco (CA)

Remote
USD 83,000 - 165,000