AI Benchmark Scientist for Biochemistry & Genomics

Obsidian

San Francisco (CA)

Remote

USD 273,000 - 382,000

Part time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will author original, executable research problems that today’s frontier models cannot solve, sourced from published papers, datasets, or open-source repositories.

The role focuses on biology-related domains with coding, including constructing prompts, defining grading criteria, and calibrating model performance.

Qualifications

  • PhD in biology, biological sciences, biochemistry, genetics, ecology, or related field.
  • Depth in at least two subdomains: ecology, biochemistry, genetics.
  • Proficient in Python, R, or similar for scientific computing.
  • Familiar with GitHub and Docker workflows.

Responsibilities

  • Write scientific prompts based on sourced input (papers, Kaggle data, open-source repos).
  • Source materials and design executable research problems for evaluation tasks.
  • Build grading criteria to define correct answers for frontier models.
  • Calibrate prompts against frontier models to ensure robust failures/successes.

Skills

Python
R
Git/GitHub
Docker
Scientific computing

Education

PhD in Biology

Tools

Docker
GitHub

Job description

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will author original, executable research problems that today’s frontier models cannot solve, sourced from published papers, datasets, or open-source repositories.

The role focuses on biology-related domains with coding, including constructing prompts, defining grading criteria, and calibrating model performance.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Biology PhD: AI Benchmark Scientist & Scientific Coder
Biology PhD: AI Benchmark Scientist & Scientific Coder

Obsidian • San Diego (CA)

On-site
USD 55,000 - 103,000
Biology PhD — AI Research Benchmark Architect
Biology PhD — AI Research Benchmark Architect

Obsidian • New York (NY)

On-site
USD 83,000 - 138,000
Biology PhD — AI Evaluation Scientist & Coder
Biology PhD — AI Evaluation Scientist & Coder

Mercor • New York (NY)

On-site
USD 55,000 - 96,000
AI Benchmark Scientist - Materials Science & Modeling (PhD)
AI Benchmark Scientist - Materials Science & Modeling (PhD)

Obsidian • New York (NY)

On-site
USD 83,000 - 138,000
AI Benchmark Scientist: Physics PhD in Scientific Computing
AI Benchmark Scientist: Physics PhD in Scientific Computing

Obsidian • Dallas (TX)

On-site
USD 9,919,000 - 14,878,000
Bio PhD AI Benchmark Architect (Part-Time, 6-Week Project)
Bio PhD AI Benchmark Architect (Part-Time, 6-Week Project)

Mercor • San Diego (CA)

On-site
USD 69,000 - 103,000
AI Benchmark Scientist for Frontier Physics Computing
AI Benchmark Scientist for Frontier Physics Computing

Mercor • Dallas (TX)

On-site
USD 83,000 - 124,000
AI Benchmark Architect for Scientific Computing
AI Benchmark Architect for Scientific Computing

Obsidian • San Francisco (CA)

Remote
USD 96,000 - 138,000
6-week engagement
Part-time 20+ hrs/week
Immediate start
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking

Obsidian • Miami (FL)

On-site
USD 83,000 - 124,000
Quantum Chemistry AI Researcher & Code Architect
Quantum Chemistry AI Researcher & Code Architect

Obsidian • San Diego (CA)

On-site
USD 83,000 - 124,000