AI Benchmark Architect for Scientific Computing

Obsidian

San Francisco (CA)

Remote

USD 96,000 - 138,000

Part time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

6-week engagement
Part-time 20+ hrs/week
Immediate start

Job summary

Mercor is seeking researchers to author AI evaluation tasks and original, executable problems for frontier models. You will source material, write prompts, and define grading criteria across subdomains with a focus on two areas in mathematics. Engagement is six weeks, part-time, with start date immediate.

Experience with Python or R for scientific computing and familiarity with Git/GitHub and Docker workflows are required to ensure reproducible runs and automated quality checks.

Qualifications

  • PhD in mathematics, applied mathematics, or a closely related field.
  • Demonstrated depth in at least two of: numerical linear algebra, computational mechanics, computational finance.
  • Working proficiency in Python or R for scientific computing.
  • Comfortable with GitHub and running code in Docker; PR workflow with automated checks.

Responsibilities

  • Source your own material: a published paper, Kaggle dataset, open-source repo, or self-designed scenario.
  • Write scientific prompts based on the input.
  • Build grading criteria that define a correct answer.
  • Calibrate against frontier models; a task ships when strong models fail it more often than succeed.

Skills

Python
R
GitHub
Docker

Education

PhD in mathematics or related field

Tools

GitHub
Docker

Job description

Mercor is seeking researchers to author AI evaluation tasks and original, executable problems for frontier models. You will source material, write prompts, and define grading criteria across subdomains with a focus on two areas in mathematics. Engagement is six weeks, part-time, with start date immediate.

Experience with Python or R for scientific computing and familiarity with Git/GitHub and Docker workflows are required to ensure reproducible runs and automated quality checks.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Scientist: Math PhD for Frontier Benchmarks
AI Evaluation Scientist: Math PhD for Frontier Benchmarks

Mercor • San Francisco (CA)

On-site
USD 11,021,000 - 13,776,000
AI Evaluation Scientist Math PhD Frontier Model Benchmark
AI Evaluation Scientist Math PhD Frontier Model Benchmark

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 179,000
AI Benchmark Scientist - Materials Science & Modeling (PhD)
AI Benchmark Scientist - Materials Science & Modeling (PhD)

Obsidian • New York (NY)

On-site
USD 83,000 - 138,000
AI Benchmark Scientist for Biochemistry & Genomics
AI Benchmark Scientist for Biochemistry & Genomics

Obsidian • San Francisco (CA)

Remote
USD 273,000 - 382,000
AI Evaluation Scientist - Math PhD (6-Week, Part-Time)
AI Evaluation Scientist - Math PhD (6-Week, Part-Time)

Mercor • Miami (FL)

On-site
USD 83,000 - 138,000
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking

Obsidian • Miami (FL)

On-site
USD 83,000 - 124,000
Bio PhD AI Benchmark Architect (Part-Time, 6-Week Project)
Bio PhD AI Benchmark Architect (Part-Time, 6-Week Project)

Mercor • San Diego (CA)

On-site
USD 69,000 - 103,000
Quantum Computing Research Engineer for AI Benchmarks
Quantum Computing Research Engineer for AI Benchmarks

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 165,000
AI Benchmark Scientist — Quantum & Computational Chemistry
AI Benchmark Scientist — Quantum & Computational Chemistry

Mercor • United States

Remote
USD 83,000 - 124,000
AI Benchmark Designer: Computational Statistics
AI Benchmark Designer: Computational Statistics

Mercor • San Francisco (CA)

On-site
USD 83,000 - 165,000