Remote AI Benchmark Scientist — Math PhD

1000scholars

United States

Remote

USD 328,000 - 546,000

Part time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Mercor is seeking PhD or Masters scientists to author AI evaluation tasks for a new Sci Code benchmark. Remote, independent contractor with a six-week, part-time commitment (20+ hours/week).

You will design problems in mathematics and related domains, source data, and establish grading criteria while calibrating against frontier models. Prior depth in numerical linear algebra, computational mechanics, or computational finance, plus Python or R proficiency, is required.

Qualifications

  • PhD in mathematics, applied mathematics, computational mathematics, or closely related field.
  • Depth in at least two subdomains: numerical linear algebra, computational mechanics, computational finance.
  • Proficiency in Python or R for scientific computing.
  • Familiarity with Git/GitHub and running code in Docker with PR workflow.

Responsibilities

  • Source your own material: published paper, Kaggle dataset, open-source repo, or self-designed scenario.
  • Write scientific prompts based on the input.
  • Build the grading criteria that define a correct answer.
  • Calibrate against frontier models so the task fails stronger than it succeeds.

Skills

Python
R
GitHub
Docker

Education

PhD in Mathematics
MS in related field

Tools

Docker
GitHub

Job description

Mercor is seeking PhD or Masters scientists to author AI evaluation tasks for a new Sci Code benchmark. Remote, independent contractor with a six-week, part-time commitment (20+ hours/week).

You will design problems in mathematics and related domains, source data, and establish grading criteria while calibrating against frontier models. Prior depth in numerical linear algebra, computational mechanics, or computational finance, plus Python or R proficiency, is required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Scientist: Math PhD for Frontier Benchmarks
AI Evaluation Scientist: Math PhD for Frontier Benchmarks

Mercor • San Francisco (CA)

On-site
USD 11,021,000 - 13,776,000
Remote Materials Science AI Benchmark Scientist (PhD)
Remote Materials Science AI Benchmark Scientist (PhD)

1000scholars • United States

Remote
USD 11,572,000 - 14,878,000
Weekly payments via Stripe or Wise
Fully remote
AI Evaluation Scientist - Math PhD (6-Week, Part-Time)
AI Evaluation Scientist - Math PhD (6-Week, Part-Time)

Mercor • Miami (FL)

On-site
USD 83,000 - 138,000
AI Benchmark Designer: Math PhD & Trainer
AI Benchmark Designer: Math PhD & Trainer

Obsidian • Los Angeles (CA)

On-site
USD 83,000 - 124,000
AI Benchmark Architect: Computational Mathematician
AI Benchmark Architect: Computational Mathematician

Mercor • San Francisco (CA)

Remote
USD 83,000 - 165,000
Remote Materials Science AI Benchmark Engineer
Remote Materials Science AI Benchmark Engineer

Weekday AI • United States

Remote
USD 83,000 - 110,000
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking

Obsidian • Miami (FL)

On-site
USD 83,000 - 124,000
Quantum AI Benchmark Scientist
Quantum AI Benchmark Scientist

Obsidian • New York (NY)

Remote
USD 83,000 - 138,000
AI Evaluation Scientist (Math PhD) & Trainer
AI Evaluation Scientist (Math PhD) & Trainer

Mercor • Los Angeles (CA)

On-site
USD 110,000 - 165,000
Remote Physics AI Benchmark Engineer (PhD)
Remote Physics AI Benchmark Engineer (PhD)

Weekday AI • United States

Remote
USD 83,000 - 117,000