AI Evaluation Scientist: Math PhD for Frontier Benchmarks

Mercor

San Francisco (CA)

On-site

USD 11,021,000 - 13,776,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, partnering with leading AI labs. You will author original, executable research problems that frontier models cannot solve.

Engagement lasts 6 weeks, part-time with 20+ hours per week, start date immediate. Required: PhD in mathematics or related field, depth in two subdomains, Python for scientific computing, and experience with GitHub and Docker.

Qualifications

  • PhD in mathematics, applied mathematics, computational mathematics, or a closely related field.
  • Demonstrated depth in at least two of: numerical linear algebra, computational mechanics, computational finance.
  • Working proficiency in Python for scientific computing.
  • Comfortable with GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design.
  • Write scientific prompts based on the input.
  • Build the grading criteria that define a correct answer.
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed.

Skills

Python for scientific computing

Education

PhD in mathematics / applied mathematics / computational mathematics

Tools

GitHub
Docker

Job description

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, partnering with leading AI labs. You will author original, executable research problems that frontier models cannot solve.

Engagement lasts 6 weeks, part-time with 20+ hours per week, start date immediate. Required: PhD in mathematics or related field, depth in two subdomains, Python for scientific computing, and experience with GitHub and Docker.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Scientist Math PhD Frontier Model Benchmark
AI Evaluation Scientist Math PhD Frontier Model Benchmark

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 179,000
AI Evaluation Scientist - Math PhD (6-Week, Part-Time)
AI Evaluation Scientist - Math PhD (6-Week, Part-Time)

Mercor • Miami (FL)

On-site
USD 83,000 - 138,000
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking

Obsidian • Miami (FL)

On-site
USD 83,000 - 124,000
AI Benchmark Designer: Math PhD & Trainer
AI Benchmark Designer: Math PhD & Trainer

Obsidian • Los Angeles (CA)

On-site
USD 83,000 - 124,000
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 179,000
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Mercor • San Francisco (CA)

On-site
USD 11,021,000 - 13,776,000
AI Evaluation Scientist (Math PhD) & Trainer
AI Evaluation Scientist (Math PhD) & Trainer

Mercor • Los Angeles (CA)

On-site
USD 110,000 - 165,000
Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Mercor • Miami (FL)

On-site
USD 83,000 - 138,000
Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Obsidian • Miami (FL)

On-site
USD 83,000 - 124,000
AI Benchmark Scientist for Frontier Physics Computing
AI Benchmark Scientist for Frontier Physics Computing

Mercor • Dallas (TX)

On-site
USD 83,000 - 124,000