Mathematics PhD - AI Evaluation Expert - AI Trainer

Mercor

Miami (FL)

On-site

USD 83,000 - 124,000

Part time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Mercor is seeking PhD and Master’s scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will source material, craft executable problems, and build the grading criteria to measure model performance on challenging scenarios.

You will calibrate prompts against frontier models, ensuring tasks remain difficult where state-of-the-art systems struggle, across two subdomains with a coding focus.

Qualifications

  • PhD in mathematics, applied mathematics, computational mathematics, or a closely related field.
  • Depth in at least two subdomains: numerical linear algebra, computational mechanics, or computational finance.
  • Proficiency in Python or R for scientific computing.
  • Experience with Git/GitHub and running code in Docker with pull-request workflows.

Responsibilities

  • Source your own material from papers, Kaggle datasets, open-source repositories, or scenarios you design.
  • Write scientific prompts based on the input.
  • Build the grading criteria that define a correct answer.
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed.

Skills

Python
R
Git/GitHub

Education

PhD in mathematics / applied mathematics / computational mathematics

Tools

Docker

Job description

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code)

Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will author original, executable research problems that today's frontier models cannot solve.

Domains - depth required in at least two subdomains (with a coding focus)
  • Mathematics — numerical linear algebra, computational mechanics, computational finance
What you'll do
  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design
  • Write scientific prompts based on the input
  • Build the grading criteria that define a correct answer
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed
Required
  • PhD in mathematics, applied mathematics, computational mathematics, or a closely related field
  • Demonstrated depth in at least two of the following subdomains: numerical linear algebra, computational mechanics, computational finance
  • Working proficiency in Python or R for scientific computing
  • Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks
Preferred
  • Publications in peer-reviewed journals
  • Prior scientific software or research engineering experience
Engagement
  • Duration: 6 weeks
  • Commitment: part-time, 20+ hours per week
  • Start date: immediate
Process
  1. Upload your resume and application form
  2. A 25-minute conversational interview covering your background, experience, and motivations
  3. Follow up within a few days with next steps and onboarding
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Mercor • Los Angeles (CA)

On-site
USD 83,000 - 124,000
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Mercor • San Francisco (CA)

On-site
USD 96,000 - 179,000
Physics PhD - Scientific Computing Expert - AI Trainer
Physics PhD - Scientific Computing Expert - AI Trainer

Mercor • Dallas (TX)

On-site
USD 9,919,000 - 14,878,000
[Contract] Mathematics PhD Coding Experts
[Contract] Mathematics PhD Coding Experts

1000scholars • United States

Remote
USD 328,000 - 546,000
Physics PhD - Scientific Computing Expert
Physics PhD - Scientific Computing Expert

Mercor • San Francisco (CA)

On-site
USD 60,000 - 120,000
Biology PhD - Scientific Coder - AI Trainer
Biology PhD - Scientific Coder - AI Trainer

Mercor • San Diego (CA)

On-site
USD 55,000 - 103,000
AI Benchmark Designer: Math PhD & Trainer
AI Benchmark Designer: Math PhD & Trainer

Mercor • Los Angeles (CA)

On-site
USD 83,000 - 124,000
Physics PhD - Quantum Computing Expert - AI Trainer
Physics PhD - Quantum Computing Expert - AI Trainer

Mercor • New York (NY)

On-site
USD 55,000 - 110,000
AI Evaluation Scientist Math PhD Frontier Model Benchmark
AI Evaluation Scientist Math PhD Frontier Model Benchmark

Mercor • San Francisco (CA)

On-site
USD 96,000 - 179,000
Materials Science PhD - AI Evaluator
Materials Science PhD - AI Evaluator

Mercor • New York (NY)

On-site
USD 83,000 - 138,000