AI Benchmark Designer: Computational Statistics

Mercor

San Francisco (CA)

On-site

USD 83,000 - 165,000

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor is building a large-scale benchmark to test how well AI systems solve advanced scientific problems. You will design graduate-level computational challenges that use real software to run simulations, interpret results, and design experiments.

The role emphasizes puzzle-design thinking over brute-force computation and requires strong Python skills. You will craft problems leveraging specialized packages in R/Python and work with Linux-based compute sandboxes, contributing to a challenging

Qualifications

  • Graduate-level training in statistics, applied mathematics, or a closely related quantitative field.
  • Proven proficiency with at least one specialized software package, demonstrated via research or professional work.
  • Strong Python skills—writing problem setups, oracle functions, and solution validators.
  • Ability to work independently and refine problem designs based on feedback.
  • Comfortable in a Linux/terminal environment with remote compute sandboxes.
  • Available for 15–20 hours per week.

Responsibilities

  • Design original problems requiring skilled use of statistical, mathematical, or scientific software.
  • Create and test problems against AI models and refine difficulty.
  • Plan problems that require multi-step workflows and strategic experimentation.

Skills

Strong Python skills
Independent work
Linux/terminal proficiency
Part-time availability 15–20 hours

Education

MS or PhD preferred; PhD preferred, or MS with 10+ years of experience

Tools

rstan
cmdstanr
brms
lavaan
deSolve
KFAS
lme4
mgcv
spatstat
mclust

Job description

Mercor is building a large-scale benchmark to test how well AI systems solve advanced scientific problems. You will design graduate-level computational challenges that use real software to run simulations, interpret results, and design experiments.

The role emphasizes puzzle-design thinking over brute-force computation and requires strong Python skills. You will craft problems leveraging specialized packages in R/Python and work with Linux-based compute sandboxes, contributing to a challenging

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Benchmark Designer: Advanced Statistical Challenges
AI Benchmark Designer: Advanced Statistical Challenges

Mercor • Los Angeles (CA)

On-site
USD 120,000 - 180,000
AI Benchmark Designer: Computational Statistics Expert
AI Benchmark Designer: Computational Statistics Expert

Mercor • Seattle (WA)

On-site
USD 60,000 - 120,000
Benchmark Designer: Computational Statistics & Applied Math
Benchmark Designer: Computational Statistics & Applied Math

Mercor • New York (NY)

On-site
USD 60,000 - 110,000
Flexible hours
Challenging bench-work
Computational Structural Engineer & AI Benchmark Designer
Computational Structural Engineer & AI Benchmark Designer

Mercor • New York (NY)

Hybrid
USD 110,000 - 170,000
Computational Statistics Benchmark Designer
Computational Statistics Benchmark Designer

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 150,000
Computational Physicist & AI Benchmark Designer
Computational Physicist & AI Benchmark Designer

Mercor • Austin (TX)

On-site
USD 83,000 - 165,000
Cosmology AI Benchmark Designer
Cosmology AI Benchmark Designer

Mercor • New York (NY)

Remote
USD 60,000 - 90,000
Bioinformatics Genomics Puzzle Designer for AI Benchmarks
Bioinformatics Genomics Puzzle Designer for AI Benchmarks

Mercor • New York (NY)

On-site
USD 69,000 - 110,000
Seismology & Geophysics AI Benchmark Designer
Seismology & Geophysics AI Benchmark Designer

Mercor • United States

Remote
USD 80,000 - 120,000
Astrophysics Benchmark Designer for AI Research
Astrophysics Benchmark Designer for AI Research

Mercor • Houston (TX)

On-site
USD 60,000 - 90,000