Mathematics PhD - AI Evaluation Expert

Mercor

San Francisco (CA)

On-site

USD 11,021,000 - 13,776,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, partnering with leading AI labs. You will author original, executable research problems that frontier models cannot solve.

Engagement lasts 6 weeks, part-time with 20+ hours per week, start date immediate. Required: PhD in mathematics or related field, depth in two subdomains, Python for scientific computing, and experience with GitHub and Docker.

Qualifications

  • PhD in mathematics, applied mathematics, computational mathematics, or a closely related field.
  • Demonstrated depth in at least two of: numerical linear algebra, computational mechanics, computational finance.
  • Working proficiency in Python for scientific computing.
  • Comfortable with GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design.
  • Write scientific prompts based on the input.
  • Build the grading criteria that define a correct answer.
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed.

Skills

Python for scientific computing

Education

PhD in mathematics / applied mathematics / computational mathematics

Tools

GitHub
Docker

Job description

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code)

Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will author original, executable research problems that today's frontier models cannot solve.

Domains - depth required in at least two subdomains (with a coding focus)

  • Mathematics — numerical linear algebra, computational mechanics, computational finance

What you'll do

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design

  • Write scientific prompts based on the input

  • Build the grading criteria that define a correct answer

  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed

Required

  • PhD in mathematics, applied mathematics, computational mathematics, or a closely related field

  • Demonstrated depth in at least two of the following subdomains: numerical linear algebra, computational mechanics, computational finance

  • Working proficiency in Python for scientific computing

  • Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks

Preferred

  • Publications in peer-reviewed journals

  • Prior scientific software or research engineering experience

Engagement

  • Duration: 6 weeks

  • Commitment: part-time, 20+ hours per week

  • Start date: immediate

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Mercor • Miami (FL)

On-site
USD 83,000 - 138,000
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 179,000
Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Obsidian • Miami (FL)

On-site
USD 83,000 - 124,000
AI Evaluation Scientist: Math PhD for Frontier Benchmarks
AI Evaluation Scientist: Math PhD for Frontier Benchmarks

Mercor • San Francisco (CA)

On-site
USD 11,021,000 - 13,776,000
Biology PhD - Scientific Coder - AI Trainer
Biology PhD - Scientific Coder - AI Trainer

Obsidian • San Diego (CA)

On-site
USD 55,000 - 103,000
Biology PhD - Scientific Coder - AI Trainer
Biology PhD - Scientific Coder - AI Trainer

Mercor • San Diego (CA)

On-site
USD 69,000 - 103,000
AI Evaluation Scientist - Math PhD (6-Week, Part-Time)
AI Evaluation Scientist - Math PhD (6-Week, Part-Time)

Mercor • Miami (FL)

On-site
USD 83,000 - 138,000
Physics PhD - Quantum Computing Expert - AI Trainer
Physics PhD - Quantum Computing Expert - AI Trainer

Mercor • New York (NY)

On-site
USD 83,000 - 124,000
Physics PhD - Quantum Computing Expert - AI Trainer
Physics PhD - Quantum Computing Expert - AI Trainer

Obsidian • New York (NY)

On-site
USD 55,000 - 110,000
Materials Science PhD - AI Evaluator
Materials Science PhD - AI Evaluator

Obsidian • New York (NY)

On-site
USD 83,000 - 138,000