Mathematics PhD - AI Evaluation Expert

Obsidian

Toronto

On-site

CAD 83,000 - 124,000

Part time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

6-week engagement
Part-time 20+ hrs/week
Immediate start

Job summary

Mercor is seeking PhD and Master's scientists to author AI evaluation tasks for Sci Code. You will source material, craft executable prompts, and design rigorous grading criteria to challenge frontier models.

This 6-week, part-time engagement emphasizes Python or R for scientific computing and requires Docker and GitHub workflow experience. Ideal candidates include published researchers or those with prior scientific software experience.

Qualifications

  • PhD in mathematics, applied mathematics, computational mathematics, or closely related field.
  • Demonstrated depth in at least two of: numerical linear algebra, computational mechanics, computational finance.
  • Working proficiency in Python or R for scientific computing.
  • Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design.
  • Write scientific prompts based on the input.
  • Build the grading criteria that define a correct answer.
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed.

Skills

Python
R
Git
Docker

Education

PhD in mathematics or related field

Tools

GitHub
Docker

Job description

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code)

Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will author original, executable research problems that today's frontier models cannot solve.

Domains - depth required in at least two subdomains (with a coding focus)
  • Mathematics — numerical linear algebra, computational mechanics, computational finance
What you'll do
  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design
  • Write scientific prompts based on the input
  • Build the grading criteria that define a correct answer
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed
Required
  • PhD in mathematics, applied mathematics, computational mathematics, or a closely related field
  • Demonstrated depth in at least two of the following subdomains: numerical linear algebra, computational mechanics, computational finance
  • Working proficiency in Python or R for scientific computing
  • Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks
Preferred
  • Publications in peer-reviewed journals
  • Prior scientific software or research engineering experience
Engagement
  • Duration: 6 weeks
  • Commitment: part-time, 20+ hours per week
  • Start date: immediate
Process
  1. Upload your resume and application form
  2. A 25-minute conversational interview covering your background, experience, and motivations
  3. Follow up within a few days with next steps and onboarding
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Mercor • Toronto

On-site
CAD 83,000 - 124,000
Biology PhD - Scientific Coder
Biology PhD - Scientific Coder

Obsidian • Toronto

On-site
CAD 34,000 - 62,000
Data Science Expert - AI Evaluation
Data Science Expert - AI Evaluation

Mercor • Toronto

On-site
CAD 120,000 - 170,000
AI Math Researcher - Question Author & Reviewer (Remote)
AI Math Researcher - Question Author & Reviewer (Remote)

Mercor • Toronto

On-site
CAD 60,000 - 96,000
Data Science Expert - Evaluation Specialist
Data Science Expert - Evaluation Specialist

Obsidian • Toronto

Hybrid
CAD 90,000 - 140,000
AI Research Scientist - PhD
AI Research Scientist - PhD

Mercor • Toronto

On-site
CAD 76,000 - 124,000
Biology AI Sci Code Architect for Benchmarks
Biology AI Sci Code Architect for Benchmarks

Obsidian • Toronto

On-site
CAD 34,000 - 62,000
AI Evaluation Scientist
AI Evaluation Scientist

Mercor • Toronto

On-site
CAD 120,000 - 180,000
Postdoctoral Research Scientist – STEM (Part-Time | $55 –$70/hr)
Postdoctoral Research Scientist – STEM (Part-Time | $55 –$70/hr)

Call For Referral • Canada

Remote
Executive Sales Representative
Executive Sales Representative

Mercor • Canada

Remote