Remote GenAI Benchmark Architect — Data Science

Mercor

New York (NY)

On-site

USD 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor is seeking experienced data scientists and quantitative analysts to join a leading AI lab's ground-truth evaluation team. You will design complex data-analysis tasks, run notebooks in Jupyter or Colab, and report clear findings to guide frontier-model research.

The role is a full-time W-2 position with Cincinnatus LLC, fully remote within the United States, and approximately 35 hours per week. You’ll work closely with researchers to ensure rigorous, reproducible evaluations.

Qualifications

  • MSc or PhD in statistics, data science, or another quantitative STEM field, or equivalent research experience.
  • 1+ years in a research, research-engineering, or heavy data-analysis role.
  • Strong data-cleaning, correlation analysis, hypothesis testing, and interpretation of results.
  • Proficiency with Jupyter Notebooks or Google Colab for analysis and reporting.
  • Working proficiency in Python (pandas, NumPy, or similar) and Git.
  • Strong ability to communicate analytical findings in writing for decision-makers.
  • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.
  • Attention to detail, creativity in task design, and ability to work independently on open-ended problems.
  • Ability to engage reliably for approximately 35 hours per week.

Responsibilities

  • Design tasks: Create realistic data-analysis challenges based on daily analytical work.
  • Author notebooks in Jupyter or Colab with clear, reproducible reference analyses.
  • Compare methods with fair evaluations and a supported recommendation.
  • Evaluate models to verify statistics and conclusions.
  • Collaborate with researchers to keep evaluations consistent and accurate.

Skills

Data cleaning
Statistical correlation
Hypothesis testing
Reporting writing
Written communication
Independent work
35 hours/week

Education

MSc/PhD in statistics or data science

Tools

Jupyter/Colab
Python (pandas/NumPy)
Git

Job description

Mercor is seeking experienced data scientists and quantitative analysts to join a leading AI lab's ground-truth evaluation team. You will design complex data-analysis tasks, run notebooks in Jupyter or Colab, and report clear findings to guide frontier-model research.

The role is a full-time W-2 position with Cincinnatus LLC, fully remote within the United States, and approximately 35 hours per week. You’ll work closely with researchers to ensure rigorous, reproducible evaluations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Quantitative Analyst for GenAI Benchmarking
Remote Quantitative Analyst for GenAI Benchmarking

Mercor • New York (NY)

Remote
USD 100,000 - 180,000
GenAI Benchmark Research Scientist - Remote (35h/wk)
GenAI Benchmark Research Scientist - Remote (35h/wk)

Mercor • San Francisco (CA)

Remote
USD 120,000 - 180,000
Remote AI Benchmark Test Engineer
Remote AI Benchmark Test Engineer

Mercor • New York (NY)

Remote
USD 85,000 - 120,000
GenAI Benchmark Research Scientist (Remote, 35h/wk)
GenAI Benchmark Research Scientist (Remote, 35h/wk)

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 160,000
GenAI Model Evaluation Engineer — Remote, 35h/wk
GenAI Model Evaluation Engineer — Remote, 35h/wk

Dorado • United States

Remote
USD 120,000 - 180,000
Remote Quantitative Analyst for AI Benchmarking
Remote Quantitative Analyst for AI Benchmarking

Mercor • New York (NY)

Remote
USD 90,000 - 130,000
Senior Software Domain Expert — GenAI QA & Benchmarks
Senior Software Domain Expert — GenAI QA & Benchmarks

Mercor • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
GenAI Benchmark Research Scientist (Remote, Part-Time)
GenAI Benchmark Research Scientist (Remote, Part-Time)

Obsidian • San Francisco (CA)

Remote
USD 120,000 - 160,000
Data Science Expert - AI/ML
Data Science Expert - AI/ML

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 160,000
GenAI Benchmark Designer - Remote Research (35h/wk)
GenAI Benchmark Designer - Remote Research (35h/wk)

Dorado • United States

Remote
USD 105,000 - 150,000