Physics PhD: AI Benchmarking Scientist for Frontier Models

Obsidian

San Francisco (CA)

On-site

USD 60,000 - 120,000

Part time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will design executable research problems that frontier models currently cannot solve and author prompts based on your sourced material.

Responsibilities include building grading criteria and calibrating against frontier models, with a focus on two or more physics subdomains and strong Python or R computing skills.

Qualifications

  • PhD in physics or closely related field.
  • Depth in at least two subdomains: condensed matter, optics, quantum information/computing, computational physics, astrophysics, particle physics.
  • Proficiency in Python, R, or similar for scientific computing.
  • Familiarity with GitHub workflows and Docker for running code.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repo, or a scenario you design
  • Write scientific prompts based on the input
  • Build the grading criteria that define a correct answer
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed

Skills

Python
R

Education

PhD in physics

Tools

GitHub
Docker

Job description

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will design executable research problems that frontier models currently cannot solve and author prompts based on your sourced material.

Responsibilities include building grading criteria and calibrating against frontier models, with a focus on two or more physics subdomains and strong Python or R computing skills.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Benchmark Scientist for Frontier Physics Computing
AI Benchmark Scientist for Frontier Physics Computing

Mercor • Dallas (TX)

On-site
USD 83,000 - 124,000
AI Evaluation Scientist: Math PhD for Frontier Benchmarks
AI Evaluation Scientist: Math PhD for Frontier Benchmarks

Mercor • San Francisco (CA)

On-site
USD 11,021,000 - 13,776,000
AI Benchmark Scientist: Physics PhD in Scientific Computing
AI Benchmark Scientist: Physics PhD in Scientific Computing

Obsidian • Dallas (TX)

On-site
USD 9,919,000 - 14,878,000
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking

Obsidian • Miami (FL)

On-site
USD 83,000 - 124,000
AI Evaluation Scientist Math PhD Frontier Model Benchmark
AI Evaluation Scientist Math PhD Frontier Model Benchmark

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 179,000
PhD Quantum Computing Scientist — AI Benchmarking
PhD Quantum Computing Scientist — AI Benchmarking

Obsidian • New York (NY)

On-site
USD 55,000 - 110,000
Quantum Chemistry AI Benchmark Engineer (PhD, Python)
Quantum Chemistry AI Benchmark Engineer (PhD, Python)

Mercor • San Diego (CA)

On-site
USD 9,919,000 - 14,878,000
Quantum Computing Scientist for AI Benchmarks (PhD)
Quantum Computing Scientist for AI Benchmarks (PhD)

Mercor • San Francisco (CA)

On-site
USD 124,000 - 165,000
Physics PhD - Scientific Computing Expert - AI Trainer
Physics PhD - Scientific Computing Expert - AI Trainer

Mercor • Dallas (TX)

On-site
USD 83,000 - 124,000
Physics PhD - Scientific Computing Expert - AI Trainer
Physics PhD - Scientific Computing Expert - AI Trainer

Obsidian • Dallas (TX)

On-site
USD 9,919,000 - 14,878,000