Quantum Computing AI Benchmark Designer

Mercor

Greater London

On-site

GBP 273,000 - 455,000

Part time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks for Sci Code. You will author original, executable research problems that today's frontier models cannot solve.

Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will craft prompts and define grading criteria, calibrating against frontier models to ensure robust evaluation results.

Qualifications

  • PhD in physics, applied physics, or closely related field.
  • Depth in at least two subdomains: condensed matter, optics, quantum information/computing, etc.
  • Proficiency in Python, R, or similar for scientific computing.
  • Comfortable with Git/GitHub and Docker, PR workflow with automated checks.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design
  • Write scientific prompts based on the input
  • Build the grading criteria that define a correct answer
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed

Skills

Python
R
Scientific computing
Git/GitHub
Docker
Research methodology

Education

PhD in physics or related field
Master's in physics or closely related field

Tools

Docker
Git
Jupyter

Job description

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks for Sci Code. You will author original, executable research problems that today's frontier models cannot solve.

Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will craft prompts and define grading criteria, calibrating against frontier models to ensure robust evaluation results.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Quantum Computing Scientist for AI Benchmarking
Quantum Computing Scientist for AI Benchmarking

Obsidian • Greater London

On-site
GBP 55,000 - 96,000
6-week engagement
Part-time 20+ hours/week
Immediate start
AI Benchmark Scientist: Biochemistry & Genetics
AI Benchmark Scientist: Biochemistry & Genetics

Obsidian • Greater London

Remote
GBP 40,000 - 50,000
AI Evaluation Scientist, Math PhD — Frontier Benchmark
AI Evaluation Scientist, Math PhD — Frontier Benchmark

Mercor • Greater London

On-site
GBP 69,000 - 124,000
Physics AI Benchmark Scientist: Prompt Design & Evaluation
Physics AI Benchmark Scientist: Prompt Design & Evaluation

Mercor • Greater London

Remote
GBP 11,000 - 18,000
AI Benchmark Scientist: Biology PhD Coder
AI Benchmark Scientist: Biology PhD Coder

Mercor • Greater London

On-site
GBP 55,000 - 83,000
Biology PhD AI Benchmark Coder
Biology PhD AI Benchmark Coder

Obsidian • Greater London

On-site
GBP 4,959,000 - 6,612,000
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Mercor • Greater London

On-site
GBP 69,000 - 124,000
Physics PhD - Quantum Computing Expert
Physics PhD - Quantum Computing Expert

Obsidian • Greater London

On-site
GBP 55,000 - 96,000
6-week engagement
Part-time 20+ hours/week
Immediate start
Physics PhD - Quantum Computing Expert
Physics PhD - Quantum Computing Expert

Mercor • Greater London

On-site
GBP 273,000 - 455,000
Computational Chemistry Benchmark Designer (AI)
Computational Chemistry Benchmark Designer (AI)

Mercor • Greater London

On-site
GBP 55,000 - 75,000