AI Evaluation Scientist - Math PhD (6-Week, Part-Time)

Mercor

Miami (FL)

On-site

USD 83,000 - 138,000

Part time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will develop original, executable research problems that frontier models cannot solve.

Expect depth in multiple subdomains, such as numerical linear algebra, computational mechanics, or computational finance, with strong Python and Docker/GitHub fluency to ensure high-quality, reproducible runs within a PR workflow.

Qualifications

  • PhD in mathematics, applied mathematics, computational mathematics, or closely related field.
  • Demonstrated depth in at least two of: numerical linear algebra, computational mechanics, computational finance.
  • Working proficiency in Python for scientific computing.
  • Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design.
  • Write scientific prompts based on the input.
  • Build the grading criteria that define a correct answer.
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed.

Skills

Python for scientific computing
Git/GitHub proficiency
Docker workflows

Education

PhD in mathematics, applied mathematics, computational mathematics, or closely related field
Master's degree in a related field

Tools

Docker
GitHub
Kaggle

Job description

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will develop original, executable research problems that frontier models cannot solve.

Expect depth in multiple subdomains, such as numerical linear algebra, computational mechanics, or computational finance, with strong Python and Docker/GitHub fluency to ensure high-quality, reproducible runs within a PR workflow.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Scientist: Math PhD for Frontier Benchmarks
AI Evaluation Scientist: Math PhD for Frontier Benchmarks

Mercor • San Francisco (CA)

On-site
USD 11,021,000 - 13,776,000
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Mercor • San Francisco (CA)

On-site
USD 11,021,000 - 13,776,000
Mathematics PhD - AI Evaluation Expert - AI Trainer
Mathematics PhD - AI Evaluation Expert - AI Trainer

Mercor • Miami (FL)

On-site
USD 83,000 - 138,000
Bio PhD AI Benchmark Architect (Part-Time, 6-Week Project)
Bio PhD AI Benchmark Architect (Part-Time, 6-Week Project)

Mercor • San Diego (CA)

On-site
USD 69,000 - 103,000
Quantum Computing AI Benchmark Scientist (PhD) — Part-Time
Quantum Computing AI Benchmark Scientist (PhD) — Part-Time

Mercor • New York (NY)

On-site
USD 83,000 - 124,000
Quantum Computing Scientist for AI Benchmarks (PhD)
Quantum Computing Scientist for AI Benchmarks (PhD)

Mercor • San Francisco (CA)

On-site
USD 124,000 - 165,000
Quantum Chemistry AI Benchmark Engineer (PhD, Python)
Quantum Chemistry AI Benchmark Engineer (PhD, Python)

Mercor • San Diego (CA)

On-site
USD 9,919,000 - 14,878,000
Biology PhD - Scientific Coder - AI Trainer
Biology PhD - Scientific Coder - AI Trainer

Mercor • San Diego (CA)

On-site
USD 69,000 - 103,000
Physics PhD - Quantum Computing Expert - AI Trainer
Physics PhD - Quantum Computing Expert - AI Trainer

Mercor • New York (NY)

On-site
USD 83,000 - 124,000
Physics PhD - Quantum Computing Expert
Physics PhD - Quantum Computing Expert

Mercor • San Francisco (CA)

On-site
USD 124,000 - 165,000