AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking

Obsidian

Miami (FL)

On-site

USD 83,000 - 124,000

Part time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor is seeking PhD and Master’s scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will source material, craft executable problems, and build the grading criteria to measure model performance on challenging scenarios.

You will calibrate prompts against frontier models, ensuring tasks remain difficult where state-of-the-art systems struggle, across two subdomains with a coding focus.

Qualifications

  • PhD in mathematics, applied mathematics, computational mathematics, or a closely related field.
  • Depth in at least two subdomains: numerical linear algebra, computational mechanics, or computational finance.
  • Proficiency in Python or R for scientific computing.
  • Experience with Git/GitHub and running code in Docker with pull-request workflows.

Responsibilities

  • Source your own material from papers, Kaggle datasets, open-source repositories, or scenarios you design.
  • Write scientific prompts based on the input.
  • Build the grading criteria that define a correct answer.
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed.

Skills

Python
R
Git/GitHub

Education

PhD in mathematics / applied mathematics / computational mathematics

Tools

Docker

Job description

Mercor is seeking PhD and Master’s scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will source material, craft executable problems, and build the grading criteria to measure model performance on challenging scenarios.

You will calibrate prompts against frontier models, ensuring tasks remain difficult where state-of-the-art systems struggle, across two subdomains with a coding focus.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Scientist: Math PhD for Frontier Benchmarks
AI Evaluation Scientist: Math PhD for Frontier Benchmarks

Mercor • San Francisco (CA)

On-site
USD 11,021,000 - 13,776,000
AI Evaluation Scientist - Math PhD (6-Week, Part-Time)
AI Evaluation Scientist - Math PhD (6-Week, Part-Time)

Mercor • Miami (FL)

On-site
USD 83,000 - 138,000
Quantum Computing Scientist for AI Benchmarks (PhD)
Quantum Computing Scientist for AI Benchmarks (PhD)

Mercor • San Francisco (CA)

On-site
USD 124,000 - 165,000
PhD Quantum Computing Scientist — AI Benchmarking
PhD Quantum Computing Scientist — AI Benchmarking

Obsidian • New York (NY)

On-site
USD 55,000 - 110,000
Quantum Computing AI Benchmark Scientist (PhD) — Part-Time
Quantum Computing AI Benchmark Scientist (PhD) — Part-Time

Mercor • New York (NY)

On-site
USD 83,000 - 124,000
Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Obsidian • Miami (FL)

On-site
USD 83,000 - 124,000
Quantum Computing Research Engineer for AI Benchmarks
Quantum Computing Research Engineer for AI Benchmarks

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 165,000
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Mercor • San Francisco (CA)

On-site
USD 11,021,000 - 13,776,000
Quantum Chemistry AI Benchmark Engineer (PhD, Python)
Quantum Chemistry AI Benchmark Engineer (PhD, Python)

Mercor • San Diego (CA)

On-site
USD 9,919,000 - 14,878,000
Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Mercor • Miami (FL)

On-site
USD 83,000 - 138,000