AI Benchmark Scientist: Physics PhD in Scientific Computing

Obsidian

Dallas (TX)

On-site

USD 9,919,000 - 14,878,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve.

In collaboration with leading AI labs, you will source your own material, write scientific prompts, build the grading criteria, and calibrate against frontier models so a task ships only when it withstands rigorous testing. The role is 6 weeks, part-time, with immediate start.

Qualifications

  • PhD in physics, applied physics, or closely related field.
  • Demonstrated depth in at least two of the following subdomains: condensed matter, optics, quantum information/computing, computational physics, astrophysics, particle physics.
  • Working proficiency in Python, R, or another relevant programming language for scientific computing.
  • Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design.
  • Write scientific prompts based on the input
  • Build the grading criteria that define a correct answer
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed

Skills

Python
R

Education

PhD in physics, applied physics, or closely related field

Tools

Git/GitHub
Docker

Job description

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve.

In collaboration with leading AI labs, you will source your own material, write scientific prompts, build the grading criteria, and calibrate against frontier models so a task ships only when it withstands rigorous testing. The role is 6 weeks, part-time, with immediate start.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Benchmark Scientist for Frontier Physics Computing
AI Benchmark Scientist for Frontier Physics Computing

Mercor • Dallas (TX)

On-site
USD 83,000 - 124,000
AI Benchmark Scientist - Materials Science & Modeling (PhD)
AI Benchmark Scientist - Materials Science & Modeling (PhD)

Obsidian • New York (NY)

On-site
USD 83,000 - 138,000
Quantum Computing Scientist for AI Benchmarks (PhD)
Quantum Computing Scientist for AI Benchmarks (PhD)

Mercor • San Francisco (CA)

On-site
USD 124,000 - 165,000
PhD Quantum Computing Scientist — AI Benchmarking
PhD Quantum Computing Scientist — AI Benchmarking

Obsidian • New York (NY)

On-site
USD 55,000 - 110,000
AI Benchmark Designer: Math PhD & Trainer
AI Benchmark Designer: Math PhD & Trainer

Obsidian • Los Angeles (CA)

On-site
USD 83,000 - 124,000
Quantum Computing Research Engineer for AI Benchmarks
Quantum Computing Research Engineer for AI Benchmarks

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 165,000
AI Evaluation Scientist Math PhD Frontier Model Benchmark
AI Evaluation Scientist Math PhD Frontier Model Benchmark

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 179,000
AI Benchmark Architect for Scientific Computing
AI Benchmark Architect for Scientific Computing

Obsidian • San Francisco (CA)

Remote
USD 96,000 - 138,000
6-week engagement
Part-time 20+ hrs/week
Immediate start
AI Evaluation Scientist (Math PhD) & Trainer
AI Evaluation Scientist (Math PhD) & Trainer

Mercor • Los Angeles (CA)

On-site
USD 110,000 - 165,000
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking

Obsidian • Miami (FL)

On-site
USD 83,000 - 124,000