Remote Data Scientist – GenAI Benchmark & Task Design

Obsidian

New York (NY)

On-site

USD 90,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cincinnatus LLC is seeking experienced data scientists and quantitative analysts to design ground-truth evaluation tasks for frontier AI models. The role emphasizes data cleaning, statistical analysis, and clear reporting in notebooks.

This is a full-time, fully remote US position with approximately 35 hours per week, collaborating with researchers to ensure rigorous evaluations that guide AI research decisions.

Qualifications

  • MSc/PhD in statistics, data science, or quantitative STEM.
  • 1+ years of experience in a research, research-engineering, or heavy data-analysis role.
  • Proficiency with Jupyter Notebooks or Google Colab for analysis and reporting.
  • Working knowledge of Python (pandas, NumPy, or similar) and Git.

Responsibilities

  • Design ground-truth evaluation tasks simulating real research work.
  • Author notebooks in Jupyter or Colab, producing clear, reproducible reference analyses.
  • Compare methods with a focus on fair evaluation and clear recommendations.
  • Evaluate how models handle the tasks and verify findings hold up statistically.
  • Collaborate with researchers to maintain consistency and accuracy in evaluations.

Skills

Data analysis
Statistical testing
Independent work
Scientific writing

Education

MSc or PhD in statistics/data science or quantitative STEM

Tools

Jupyter/Colab
Python
Git
NumPy

Job description

Cincinnatus LLC is seeking experienced data scientists and quantitative analysts to design ground-truth evaluation tasks for frontier AI models. The role emphasizes data cleaning, statistical analysis, and clear reporting in notebooks.

This is a full-time, fully remote US position with approximately 35 hours per week, collaborating with researchers to ensure rigorous evaluations that guide AI research decisions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Benchmark Research Scientist (Remote, 35h/wk)
GenAI Benchmark Research Scientist (Remote, 35h/wk)

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 160,000
GenAI Benchmark Research Scientist (Remote, Part-Time)
GenAI Benchmark Research Scientist (Remote, Part-Time)

Obsidian • San Francisco (CA)

Remote
USD 120,000 - 160,000
Remote Quantitative Analyst for AI Benchmarking
Remote Quantitative Analyst for AI Benchmarking

Mercor • New York (NY)

Remote
USD 90,000 - 130,000
GenAI Benchmark Designer - Remote Research (35h/wk)
GenAI Benchmark Designer - Remote Research (35h/wk)

Dorado • United States

Remote
USD 105,000 - 150,000
Remote Quantitative Analyst for GenAI Benchmarking
Remote Quantitative Analyst for GenAI Benchmarking

Mercor • New York (NY)

Remote
USD 100,000 - 180,000
Remote Data Science & Quant Analytics Expert
Remote Data Science & Quant Analytics Expert

Dorado • United States

Remote
USD 90,000 - 120,000
Remote GenAI Benchmark Architect — Data Science
Remote GenAI Benchmark Architect — Data Science

Mercor • New York (NY)

On-site
USD 120,000 - 170,000
GenAI Evaluation Scientist (Remote, 35h/wk)
GenAI Evaluation Scientist (Remote, 35h/wk)

Mercor • New York (NY)

Remote
USD 90,000 - 120,000
GenAI Benchmark Research Scientist — Remote, Part-Time
GenAI Benchmark Research Scientist — Remote, Part-Time

Obsidian • New York (NY)

Remote
USD 120,000 - 150,000
Remote QA/Test Engineer for AI Benchmarks
Remote QA/Test Engineer for AI Benchmarks

Dorado • United States

Remote
USD 90,000 - 130,000