Mathematics PhD - AI Evaluation Expert - AI Trainer

Mercor

Los Angeles (CA)

On-site

USD 110,000 - 165,000

Part time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Mercor is seeking PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will author original, executable research problems that today's frontier models cannot solve.

In Los Angeles, this part-time engagement runs 6 weeks with 20+ hours per week. You will source materials, write scientific prompts, build grading criteria, and calibrate against frontier models to ensure tasks fail strong models more often than they succeed.

Qualifications

  • PhD in mathematics, applied mathematics, computational mathematics, or a closely related field.
  • Demonstrated depth in at least two of: numerical linear algebra, computational mechanics, computational finance.
  • Working proficiency in Python or R for scientific computing; experience with Docker and GitHub workflows.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design.
  • Write scientific prompts based on the input.
  • Build the grading criteria that define a correct answer.
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed.

Skills

Python
R
Numerical linear algebra
Computational mechanics
Computational finance

Education

PhD in mathematics / applied mathematics

Tools

Git/GitHub
Docker

Job description

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code)

Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will author original, executable research problems that today's frontier models cannot solve.

Domains - depth required in at least two subdomains (with a coding focus)
  • Mathematics — numerical linear algebra, computational mechanics, computational finance
What you'll do
  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design

  • Write scientific prompts based on the input

  • Build the grading criteria that define a correct answer

  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed

Required
  • PhD in mathematics, applied mathematics, computational mathematics, or a closely related field

  • Demonstrated depth in at least two of the following subdomains: numerical linear algebra, computational mechanics, computational finance

  • Working proficiency in Python or R for scientific computing

  • Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks

Preferred
  • Publications in peer-reviewed journals

  • Prior scientific software or research engineering experience

Engagement
  • Duration: 6 weeks

  • Commitment: part-time, 20+ hours per week

  • Start date: immediate

Process
  1. A 25-minute conversational interview covering your background, experience, and motivations

  2. Follow up within a few days with next steps and onboarding

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Obsidian • Los Angeles (CA)

On-site
USD 83,000 - 124,000
Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Obsidian • Miami (FL)

On-site
USD 83,000 - 124,000
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 179,000
Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Mercor • Miami (FL)

On-site
USD 83,000 - 138,000
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Mercor • San Francisco (CA)

On-site
USD 11,021,000 - 13,776,000
Physics PhD - Scientific Computing Expert - AI Trainer
Physics PhD - Scientific Computing Expert - AI Trainer

Mercor • Dallas (TX)

On-site
USD 83,000 - 124,000
Physics PhD - Scientific Computing Expert - AI Trainer
Physics PhD - Scientific Computing Expert - AI Trainer

Obsidian • Dallas (TX)

On-site
USD 9,919,000 - 14,878,000
AI Evaluation Scientist: Math PhD for Frontier Benchmarks
AI Evaluation Scientist: Math PhD for Frontier Benchmarks

Mercor • San Francisco (CA)

On-site
USD 11,021,000 - 13,776,000
Physics PhD - Scientific Computing Expert
Physics PhD - Scientific Computing Expert

Obsidian • San Francisco (CA)

On-site
USD 60,000 - 120,000
Biology PhD - Scientific Coder - AI Trainer
Biology PhD - Scientific Coder - AI Trainer

Mercor • San Diego (CA)

On-site
USD 69,000 - 103,000