GenAI Benchmark Research Scientist (Remote, 35h/wk)

Obsidian

San Francisco (CA)

On-site

USD 100,000 - 160,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Cincinnatus LLC is recruiting experienced data scientists and quantitative analysts to design ground-truth tasks for frontier AI models. This full-time W-2 role is fully remote in the United States, approx. 35 hours per week, with tasks ranging from data cleaning to reporting in notebooks.

You will work closely with researchers to ensure rigorous analyses and clear recommendations for decision-makers. The role requires 1+ year in research-heavy analytics, strong Python (pandas/NumPy), Jupyter or

Qualifications

  • MSc or PhD in statistics, data science, or another quantitative STEM field, or equivalent practical experience in a research-heavy analytical domain.
  • 1+ years of experience in a research, research-engineering, or heavy data-analysis role.
  • Hands-on data-analysis skills: data cleaning, statistical correlation, hypothesis testing, and interpretation.
  • Proficiency with Jupyter Notebooks or Google Colab for analysis and reporting.
  • Working proficiency in Python (pandas, NumPy) and Git.
  • Strong ability to communicate analytical findings in writing for decision-makers.
  • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.
  • Perfectionist mindset with high attention to detail and independent work in open-ended problems.
  • Ability to engage reliably for approximately 35 hours per week.

Responsibilities

  • Design tasks: Create realistic data-analysis challenges — cleaning messy data, comparing methods, interpreting results.
  • Author notebooks: Work through tasks in Jupyter or Colab, producing clear, reproducible analyses.
  • Compare methods: Build tasks for fair comparisons between analytical approaches with spot checks.
  • Evaluate models: Review how models handle tasks, ensuring statistics hold up.
  • Work as a team: Compare notes with researchers to keep evaluations consistent and accurate.

Skills

Python
Pandas/NumPy
Data cleaning
Statistical analysis
Git
Notebook reporting

Education

MSc/PhD in statistics or data science
1+ years research/data-analysis experience

Tools

Jupyter
Colab

Job description

Cincinnatus LLC is recruiting experienced data scientists and quantitative analysts to design ground-truth tasks for frontier AI models. This full-time W-2 role is fully remote in the United States, approx. 35 hours per week, with tasks ranging from data cleaning to reporting in notebooks.

You will work closely with researchers to ensure rigorous analyses and clear recommendations for decision-makers. The role requires 1+ year in research-heavy analytics, strong Python (pandas/NumPy), Jupyter or

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GenAI Benchmark Research Scientist (Remote, Part-Time)
GenAI Benchmark Research Scientist (Remote, Part-Time)

Obsidian • San Francisco (CA)

Remote
USD 120,000 - 160,000
Remote Data Scientist – GenAI Benchmark & Task Design
Remote Data Scientist – GenAI Benchmark & Task Design

Obsidian • New York (NY)

On-site
USD 90,000 - 150,000
Remote Quantitative Analyst for AI Benchmarking
Remote Quantitative Analyst for AI Benchmarking

Mercor • New York (NY)

Remote
USD 90,000 - 130,000
GenAI Benchmark Research Scientist — Remote, Part-Time
GenAI Benchmark Research Scientist — Remote, Part-Time

Obsidian • New York (NY)

Remote
USD 120,000 - 150,000
GenAI Evaluation Scientist (Remote, 35h/wk)
GenAI Evaluation Scientist (Remote, 35h/wk)

Mercor • New York (NY)

Remote
USD 90,000 - 120,000
Remote Quantitative Analyst for GenAI Benchmarking
Remote Quantitative Analyst for GenAI Benchmarking

Mercor • New York (NY)

Remote
USD 100,000 - 180,000
GenAI Benchmark Research Scientist - Remote (35h/wk)
GenAI Benchmark Research Scientist - Remote (35h/wk)

Mercor • San Francisco (CA)

Remote
USD 120,000 - 180,000
Remote GenAI Benchmark Architect — Data Science
Remote GenAI Benchmark Architect — Data Science

Mercor • New York (NY)

On-site
USD 120,000 - 170,000
GenAI Vulnerability Researcher (Remote, 35h/w)
GenAI Vulnerability Researcher (Remote, 35h/w)

Obsidian • San Francisco (CA)

Remote
USD 130,000 - 160,000
Senior AI Engineering Specialist — GenAI Benchmarks
Senior AI Engineering Specialist — GenAI Benchmarks

Mercor • San Francisco (CA)

Hybrid
USD 180,000 - 240,000