AI Evaluation Scientist (Math PhD) & Trainer

Mercor

Los Angeles (CA)

On-site

USD 110,000 - 165,000

Part time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Mercor is seeking PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will author original, executable research problems that today's frontier models cannot solve.

In Los Angeles, this part-time engagement runs 6 weeks with 20+ hours per week. You will source materials, write scientific prompts, build grading criteria, and calibrate against frontier models to ensure tasks fail strong models more often than they succeed.

Qualifications

  • PhD in mathematics, applied mathematics, computational mathematics, or a closely related field.
  • Demonstrated depth in at least two of: numerical linear algebra, computational mechanics, computational finance.
  • Working proficiency in Python or R for scientific computing; experience with Docker and GitHub workflows.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design.
  • Write scientific prompts based on the input.
  • Build the grading criteria that define a correct answer.
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed.

Skills

Python
R
Numerical linear algebra
Computational mechanics
Computational finance

Education

PhD in mathematics / applied mathematics

Tools

Git/GitHub
Docker

Job description

Mercor is seeking PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will author original, executable research problems that today's frontier models cannot solve.

In Los Angeles, this part-time engagement runs 6 weeks with 20+ hours per week. You will source materials, write scientific prompts, build grading criteria, and calibrate against frontier models to ensure tasks fail strong models more often than they succeed.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking

Obsidian • Miami (FL)

On-site
USD 83,000 - 124,000
AI Evaluation Scientist - Math PhD (6-Week, Part-Time)
AI Evaluation Scientist - Math PhD (6-Week, Part-Time)

Mercor • Miami (FL)

On-site
USD 83,000 - 138,000
AI Evaluation Scientist Math PhD Frontier Model Benchmark
AI Evaluation Scientist Math PhD Frontier Model Benchmark

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 179,000
AI Evaluation Scientist: Math PhD for Frontier Benchmarks
AI Evaluation Scientist: Math PhD for Frontier Benchmarks

Mercor • San Francisco (CA)

On-site
USD 11,021,000 - 13,776,000
AI Benchmark Scientist: Physics PhD in Scientific Computing
AI Benchmark Scientist: Physics PhD in Scientific Computing

Obsidian • Dallas (TX)

On-site
USD 9,919,000 - 14,878,000
AI Benchmark Designer: Math PhD & Trainer
AI Benchmark Designer: Math PhD & Trainer

Obsidian • Los Angeles (CA)

On-site
USD 83,000 - 124,000
Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Mercor • Miami (FL)

On-site
USD 83,000 - 138,000
AI Benchmark Scientist - Materials Science & Modeling (PhD)
AI Benchmark Scientist - Materials Science & Modeling (PhD)

Obsidian • New York (NY)

On-site
USD 83,000 - 138,000
AI Benchmark Architect for Scientific Computing
AI Benchmark Architect for Scientific Computing

Obsidian • San Francisco (CA)

Remote
USD 96,000 - 138,000
6-week engagement
Part-time 20+ hrs/week
Immediate start
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Mercor • San Francisco (CA)

On-site
USD 11,021,000 - 13,776,000