6-Week Part-Time Materials Scientist for AI Benchmarks

Mercor

San Francisco (CA)

Remote

USD 150,000 - 190,000

Part time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will author original, executable research problems that today's frontier models cannot solve.

Source your own material, write prompts, and build grading criteria; you will calibrate against frontier models so a task ships only when strong models fail more often than they succeed. This six-week, part-time engagement starts immediately.

Qualifications

  • PhD in materials science, materials engineering, applied physics, chemistry, or chemical engineering.
  • Demonstrated depth in semiconductor materials and molecular modeling.
  • Working proficiency in Python, R, or another relevant programming language for scientific computing.
  • Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks.

Responsibilities

  • Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design.
  • Write scientific prompts based on the input.
  • Build the grading criteria that define a correct answer.
  • Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed.

Skills

Python
R
Git/GitHub
Docker

Education

PhD in materials science, materials engineering, applied physics, chemistry, chemical engineering

Tools

Git/GitHub
Docker

Job description

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will author original, executable research problems that today's frontier models cannot solve.

Source your own material, write prompts, and build grading criteria; you will calibrate against frontier models so a task ships only when strong models fail more often than they succeed. This six-week, part-time engagement starts immediately.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Benchmark Scientist - Materials Science & Modeling (PhD)
AI Benchmark Scientist - Materials Science & Modeling (PhD)

Obsidian • New York (NY)

On-site
USD 83,000 - 138,000
Materials Scientist: AI Benchmarking & Research Prompts
Materials Scientist: AI Benchmarking & Research Prompts

Obsidian • San Francisco (CA)

Remote
USD 8,266,000 - 12,398,000
Quantum Computing Research Engineer for AI Benchmarks
Quantum Computing Research Engineer for AI Benchmarks

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 165,000
Quantum Computing Scientist for AI Benchmarks (PhD)
Quantum Computing Scientist for AI Benchmarks (PhD)

Mercor • San Francisco (CA)

On-site
USD 124,000 - 165,000
AI Benchmark Architect: Computational Mathematician
AI Benchmark Architect: Computational Mathematician

Mercor • San Francisco (CA)

Remote
USD 83,000 - 165,000
Bio PhD AI Benchmark Architect (Part-Time, 6-Week Project)
Bio PhD AI Benchmark Architect (Part-Time, 6-Week Project)

Mercor • San Diego (CA)

On-site
USD 69,000 - 103,000
Quantum AI Benchmark Scientist
Quantum AI Benchmark Scientist

Obsidian • New York (NY)

Remote
USD 83,000 - 138,000
Biochemist for AI Benchmark Tasks & Prompt Design
Biochemist for AI Benchmark Tasks & Prompt Design

Mercor • San Francisco (CA)

Remote
USD 96,000 - 138,000
AI Benchmark Scientist: Physics PhD in Scientific Computing
AI Benchmark Scientist: Physics PhD in Scientific Computing

Obsidian • Dallas (TX)

On-site
USD 9,919,000 - 14,878,000
AI Benchmark Scientist for Frontier Physics Computing
AI Benchmark Scientist for Frontier Physics Computing

Mercor • Dallas (TX)

On-site
USD 83,000 - 124,000