AI Benchmark Problem Designer: Computational Statistics

Obsidian

Seattle (WA)

On-site

USD 120,000 - 180,000

Part time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Obsidian is seeking a Computational Statistics and Applied Mathematics Expert to design graduate-level problems for AI benchmarks. You will craft tasks requiring real scientific software and multi-step workflows that test planning, data interpretation, and experimental design.

You will test problems against state-of-the-art AI models, iterating until the target difficulty is met. This role supports remote work and part-time hours (15–20 per week), with Seattle area preference.

Qualifications

  • Graduate-level training in statistics, applied mathematics, or a closely related field.
  • Proven proficiency with at least one specialized statistical, mathematical, or scientific software package, demonstrated through research publications, open-source contributions, or professional work
  • Strong Python skills — you'll be writing problem setups, oracle functions, and solution validators
  • Ability to work independently and refine problem designs based on feedback
  • Comfortable working in a Linux/terminal environment with remote compute sandboxes
  • Available for at least 15–20 hours per week

Responsibilities

  • Create problems requiring skilled use of specialized statistical, mathematical, or scientific software packages.
  • Test problems against state-of-the-art AI models and refine them until the target difficulty is reached.
  • Design and interpret numerical results from fully defined setups.
  • Plan experiments and queries to uncover information not directly visible.
  • Work through testing loops against modern AI models and iterate accordingly.

Skills

Python programming
Statistical analysis
Applied mathematics
Problem design
Independent work

Education

MS or PhD in statistics / applied math

Tools

R
Python
Matlab/Scilab
Linux/Unix

Job description

Obsidian is seeking a Computational Statistics and Applied Mathematics Expert to design graduate-level problems for AI benchmarks. You will craft tasks requiring real scientific software and multi-step workflows that test planning, data interpretation, and experimental design.

You will test problems against state-of-the-art AI models, iterating until the target difficulty is met. This role supports remote work and part-time hours (15–20 per week), with Seattle area preference.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Benchmark Problem Designer — Applied Mathematics
AI Benchmark Problem Designer — Applied Mathematics

Obsidian • San Francisco (CA)

Remote
USD 140,000 - 210,000
Computational Statistics Benchmark Designer
Computational Statistics Benchmark Designer

Mercor • San Francisco (CA)

On-site
USD 120,000 - 150,000
Bayesian Computation & Problem Design Scientist
Bayesian Computation & Problem Design Scientist

Obsidian • San Francisco (CA)

On-site
USD 80,000 - 120,000
Senior AI Benchmark Problem Designer (Contract)
Senior AI Benchmark Problem Designer (Contract)

SmartRecruiters, Inc. • Florida City (FL)

On-site
USD 90,000 - 130,000
AI/ML Software Engineer — Benchmark Task Designer
AI/ML Software Engineer — Benchmark Task Designer

Obsidian • San Francisco (CA)

On-site
USD 68,880 - 103,320
Computational Bayesian Problem Designer
Computational Bayesian Problem Designer

Mercor • San Francisco (CA)

Hybrid
USD 70,000 - 100,000
Astrophysics Problem Designer for AI Benchmarks
Astrophysics Problem Designer for AI Benchmarks

Obsidian • Houston (TX)

On-site
USD 70,000 - 120,000
Astrophysics AI Benchmark Designer - Expert
Astrophysics AI Benchmark Designer - Expert

Mercor • New York (NY)

On-site
USD 70,000 - 110,000
Astrophysics AI Benchmark Designer (Computational)
Astrophysics AI Benchmark Designer (Computational)

Mercor • San Francisco (CA)

On-site
USD 55,000 - 83,000
Remote Math Expert for AI Benchmarking & Problem Design
Remote Math Expert for AI Benchmarking & Problem Design

Anyone AI Inc. • United States

Remote
MXN 619,000 - 991,000