Remote Quantitative Analyst for AI Benchmarking

Mercor

New York (NY)

Remote

USD 90,000 - 130,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Cincinnatus LLC in the United States seeks experienced data scientists and quantitative analysts to design ground-truth evaluation tasks for frontier models. You will create realistic data-analysis challenges, perform cleaning, compute correlations, test hypotheses, and summarize findings for researchers in a notebook.

This full-time, remote W-2 role covers approximately 35 hours per week and involves collaboration with lab researchers, task design, notebook authoring in Jupyter or Colab, and

Qualifications

  • MSc or PhD in statistics, data science, or another quantitative STEM field, or equivalent practical experience in a research-heavy analytical domain.
  • 1+ years of experience in a research, research-engineering, or heavy data-analysis role.
  • Deep hands‑on data-analysis skills: data cleaning, statistical correlation, hypothesis testing, and careful interpretation of results.
  • Proficiency with Jupyter Notebooks or Google Colab for analysis and reporting.
  • Working proficiency in Python (pandas, NumPy, or similar) and Git.
  • Strong ability to communicate analytical findings in writing for decision-makers.
  • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.
  • A perfectionist mindset: high attention to detail, creativity in task design, and the ability to work independently through ambiguous, open‑ended problems.
  • Ability to engage reliably for approximately 35 hours per week.

Responsibilities

  • Design tasks: Create realistic data-analysis challenges — cleaning messy data, comparing methods, interpreting results — based on the kind of analysis you do every day.
  • Author notebooks: Work through your own tasks in Jupyter or Colab, producing clear, reproducible reference analyses.
  • Compare methods: Build tasks that ask for a fair comparison between analytical approaches, backed by spot checks and a clear recommendation.
  • Evaluate models: Review how models handle your tasks, and check whether their statistics and conclusions actually hold up.
  • Work as a team: Compare notes with researchers and fellow experts to keep evaluations consistent and accurate.

Skills

Data analysis
Python (pandas)
Jupyter/Colab
Statistics
Communication
Git
Research experience
Independent work

Education

MSc/PhD in statistics or data science
Equivalent practical experience

Tools

Jupyter Notebook
Colab

Job description

Cincinnatus LLC in the United States seeks experienced data scientists and quantitative analysts to design ground-truth evaluation tasks for frontier models. You will create realistic data-analysis challenges, perform cleaning, compute correlations, test hypotheses, and summarize findings for researchers in a notebook.

This full-time, remote W-2 role covers approximately 35 hours per week and involves collaboration with lab researchers, task design, notebook authoring in Jupyter or Colab, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Quantitative Analyst for GenAI Benchmarking
Remote Quantitative Analyst for GenAI Benchmarking

Mercor • New York (NY)

Remote
USD 100,000 - 180,000
GenAI Benchmark Research Scientist (Remote, 35h/wk)
GenAI Benchmark Research Scientist (Remote, 35h/wk)

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 160,000
Remote Data Scientist – GenAI Benchmark & Task Design
Remote Data Scientist – GenAI Benchmark & Task Design

Obsidian • New York (NY)

On-site
USD 90,000 - 150,000
GenAI Benchmark Research Scientist (Remote, Part-Time)
GenAI Benchmark Research Scientist (Remote, Part-Time)

Obsidian • San Francisco (CA)

Remote
USD 120,000 - 160,000
Remote GenAI Benchmark Architect — Data Science
Remote GenAI Benchmark Architect — Data Science

Mercor • New York (NY)

On-site
USD 120,000 - 170,000
GenAI Benchmark Research Scientist — Remote, Part-Time
GenAI Benchmark Research Scientist — Remote, Part-Time

Obsidian • New York (NY)

Remote
USD 120,000 - 150,000
Data Science Expert - AI/ML
Data Science Expert - AI/ML

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 160,000
Remote AI Benchmark Test Engineer
Remote AI Benchmark Test Engineer

Mercor • New York (NY)

Remote
USD 85,000 - 120,000
Remote QA/Test Engineer — AI Benchmark Validation
Remote QA/Test Engineer — AI Benchmark Validation

Weekday AI (YC W21) • United States

On-site
USD 83,000 - 124,000
Senior Life Sciences Scientist - AI QA & Benchmarking
Senior Life Sciences Scientist - AI QA & Benchmarking

Obsidian • New York (NY)

Remote
USD 180,000 - 240,000