GenAI Benchmark Research Scientist (Remote, Part-Time)

Obsidian

San Francisco (CA)

Remote

USD 120,000 - 160,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Cincinnatus LLC is recruiting researchers to design and author multi-step evaluation tasks for frontier AI benchmarks. The role emphasizes translating scientific method into practical tasks, with a focus on Python-based analysis, rigorous evaluation, and clear written conclusions.

You will work remotely in the United States for approximately 35 hours per week as part of a full-time W-2 engagement, integrated with leading AI lab teams through Cincinnatus’ extended workforce model.

Qualifications

  • MSc or PhD in a STEM field, or in a computational social-science or humanities discipline, or equivalent practical experience in a research-heavy domain requiring data analysis and coding.
  • 1+ years of experience in an active research role (academia, industry, or national labs).
  • Your own research involves significant computational work: Python-based analysis, simulation, modeling, or data pipelines.
  • Strong grounding in experimental design, hypothesis testing, and rigorous evaluation of results.
  • Working familiarity with Git, IDEs, and notebook environments (Jupyter or Colab).
  • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.
  • A perfectionist mindset: high attention to detail, creativity in task design, strong written communication, and the ability to work independently through ambiguous, open-ended problems.
  • Ability to engage reliably for approximately 35 hours per week.

Responsibilities

  • Design tasks: Turn the research skills you use every day — designing studies, testing hypotheses, evaluating results — into engaging, multi-step tasks.

Skills

Python data analysis
Experimental design
Hypothesis testing
Notebook environments
Research experience
Git version control
Academic writing/communication
Independent work
35 hours/week

Education

MSc or PhD in STEM or computational social science/humanities

Tools

Git
Jupyter/Colab

Job description

Cincinnatus LLC is recruiting researchers to design and author multi-step evaluation tasks for frontier AI benchmarks. The role emphasizes translating scientific method into practical tasks, with a focus on Python-based analysis, rigorous evaluation, and clear written conclusions.

You will work remotely in the United States for approximately 35 hours per week as part of a full-time W-2 engagement, integrated with leading AI lab teams through Cincinnatus’ extended workforce model.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Benchmark Research Scientist (Remote, 35h/wk)
GenAI Benchmark Research Scientist (Remote, 35h/wk)

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 160,000
GenAI Benchmark Research Scientist — Remote, Part-Time
GenAI Benchmark Research Scientist — Remote, Part-Time

Obsidian • New York (NY)

Remote
USD 120,000 - 150,000
Remote Data Scientist – GenAI Benchmark & Task Design
Remote Data Scientist – GenAI Benchmark & Task Design

Obsidian • New York (NY)

On-site
USD 90,000 - 150,000
GenAI Evaluation Scientist (Remote, 35h/wk)
GenAI Evaluation Scientist (Remote, 35h/wk)

Mercor • New York (NY)

Remote
USD 90,000 - 120,000
GenAI Benchmark Research Scientist - Remote (35h/wk)
GenAI Benchmark Research Scientist - Remote (35h/wk)

Mercor • San Francisco (CA)

Remote
USD 120,000 - 180,000
GenAI Vulnerability Researcher (Remote, 35h/w)
GenAI Vulnerability Researcher (Remote, 35h/w)

Obsidian • San Francisco (CA)

Remote
USD 130,000 - 160,000
Remote Quantitative Analyst for AI Benchmarking
Remote Quantitative Analyst for AI Benchmarking

Mercor • New York (NY)

Remote
USD 90,000 - 130,000
Remote Quantitative Analyst for GenAI Benchmarking
Remote Quantitative Analyst for GenAI Benchmarking

Mercor • New York (NY)

Remote
USD 100,000 - 180,000
Senior AI Engineering Specialist — GenAI Benchmarks
Senior AI Engineering Specialist — GenAI Benchmarks

Mercor • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Remote GenAI Benchmark Architect — Data Science
Remote GenAI Benchmark Architect — Data Science

Mercor • New York (NY)

On-site
USD 120,000 - 170,000