AI Benchmark Architect for Scientific Computing

Mercor

Toronto

Remote

CAD 83,000 - 165,000

Part time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will craft original, executable research problems that frontier models currently struggle to solve.

Domains require depth in at least two subfields with a coding focus; Python or R for scientific computing is essential. The role involves sourcing material, writing prompts, and building robust grading criteria for model evaluation.

Qualifications

  • PhD in mathematics or closely related field required.
  • Demonstrated depth in at least two subdomains (e.g., numerical linear algebra, computational mechanics, computational finance).
  • Proficiency in Python or R for scientific computing and ability to run code in Docker with PR workflows.

Responsibilities

  • Source your own material: published paper, Kaggle dataset, open-source repo, or a scenario you design.
  • Write scientific prompts based on the input and design executable research problems.
  • Build the grading criteria defining a correct answer and calibrate against frontier models.

Skills

Python
R
Scientific computing
Git/GitHub
Docker
Mathematics depth

Education

PhD in mathematics
Master's degree in related field

Tools

GitHub
Docker

Job description

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will craft original, executable research problems that frontier models currently struggle to solve.

Domains require depth in at least two subfields with a coding focus; Python or R for scientific computing is essential. The role involves sourcing material, writing prompts, and building robust grading criteria for model evaluation.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Biology AI Sci Code Architect for Benchmarks
Biology AI Sci Code Architect for Benchmarks

Obsidian • Toronto

On-site
CAD 34,000 - 62,000
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Mercor • Toronto

On-site
CAD 83,000 - 124,000
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Obsidian • Toronto

On-site
CAD 83,000 - 124,000
6-week engagement
Part-time 20+ hrs/week
Immediate start
Biology PhD - Scientific Coder
Biology PhD - Scientific Coder

Obsidian • Toronto

On-site
CAD 34,000 - 62,000
AI Evaluation Scientist
AI Evaluation Scientist

Mercor • Toronto

On-site
CAD 120,000 - 180,000
AI Benchmark Designer: Computational Statistics
AI Benchmark Designer: Computational Statistics

Obsidian • Toronto

On-site
CAD 90,000 - 120,000
Remote AI Math Assessment Architect (PhD)
Remote AI Math Assessment Architect (PhD)

Mercor • Toronto

On-site
CAD 70,000 - 110,000
Fully remote
Data Science Quality Architect for AI Evaluation
Data Science Quality Architect for AI Evaluation

Obsidian • Toronto

Remote
CAD 110,000 - 170,000
Data Science Expert - Evaluation Specialist
Data Science Expert - Evaluation Specialist

Mercor • Toronto

On-site
CAD 120,000 - 180,000
Data Science Expert - Evaluation Specialist
Data Science Expert - Evaluation Specialist

Obsidian • Toronto

On-site
CAD 90,000 - 140,000