Mathematics PhD - AI Evaluation Expert

Obsidian

San Francisco (CA)

On-site

USD 96,000 - 179,000

Part time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve. The role emphasizes material sourcing, prompt design, and robust grading criteria against leading AI models.

Engagement spans 6 weeks, part-time at 20+ hours per week, with immediate start. You will work with Python or R for scientific computing and use GitHub and Docker in a PR-driven workflow.

Qualifications

  • PhD in mathematics, applied mathematics, or closely related field.
  • Depth in at least two subdomains: numerical linear algebra, computational mechanics or computational finance.
  • Working proficiency in Python or R for scientific computing.
  • Comfortable with Git/GitHub and running code in Docker with PR workflows.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design.
  • Write scientific prompts based on the input.
  • Build the grading criteria that define a correct answer.
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed.

Skills

Numerical linear algebra
Computational mechanics
Python or R for scientific computing

Education

PhD in mathematics or related field

Tools

Git/GitHub
Docker

Job description

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code)

Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will author original, executable research problems that today's frontier models cannot solve.

Domains - depth required in at least two subdomains (with a coding focus)
  • Mathematics — numerical linear algebra, computational mechanics, computational finance
What you'll do
  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design

  • Write scientific prompts based on the input

  • Build the grading criteria that define a correct answer

  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed

Required
  • PhD in mathematics, applied mathematics, computational mathematics, or a closely related field

  • Demonstrated depth in at least two of the following subdomains: numerical linear algebra, computational mechanics, computational finance

  • Working proficiency in Python or R for scientific computing

  • Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks

Preferred
  • Publications in peer-reviewed journals

  • Prior scientific software or research engineering experience

Engagement
  • Duration: 6 weeks

  • Commitment: part-time, 20+ hours per week

  • Start date: immediate

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Obsidian • Miami (FL)

On-site
USD 83,000 - 124,000
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Mercor • San Francisco (CA)

On-site
USD 11,021,000 - 13,776,000
Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Mercor • Miami (FL)

On-site
USD 83,000 - 138,000
AI Evaluation Scientist: Math PhD for Frontier Benchmarks
AI Evaluation Scientist: Math PhD for Frontier Benchmarks

Mercor • San Francisco (CA)

On-site
USD 11,021,000 - 13,776,000
AI Evaluation Scientist - Math PhD (6-Week, Part-Time)
AI Evaluation Scientist - Math PhD (6-Week, Part-Time)

Mercor • Miami (FL)

On-site
USD 83,000 - 138,000
AI Evaluation Scientist Math PhD Frontier Model Benchmark
AI Evaluation Scientist Math PhD Frontier Model Benchmark

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 179,000
Biology PhD - Scientific Coder - AI Trainer
Biology PhD - Scientific Coder - AI Trainer

Obsidian • San Diego (CA)

On-site
USD 55,000 - 103,000
Biology PhD - Scientific Coder - AI Trainer
Biology PhD - Scientific Coder - AI Trainer

Mercor • San Diego (CA)

On-site
USD 69,000 - 103,000
Physics PhD - Quantum Computing Expert - AI Trainer
Physics PhD - Quantum Computing Expert - AI Trainer

Mercor • New York (NY)

On-site
USD 83,000 - 124,000
Materials Science PhD - AI Evaluator
Materials Science PhD - AI Evaluator

Obsidian • New York (NY)

On-site
USD 83,000 - 138,000