GenAI Benchmark Research Scientist — Remote, Part-Time

Obsidian

New York (NY)

Remote

USD 120,000 - 150,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Cincinnatus LLC is seeking experienced ML practitioners to author complex, multi-step ML tasks for frontier models. You will design experiments, implement changes, run trainings, and analyze results to determine where models fall short.

This full-time, fully remote role within the United States requires a MSc/PhD or equivalent, 1+ years research experience, and strong Python/Git skills, with a focus on RL and LLM evaluation.

Qualifications

  • MSc or PhD in ML, CS, or equivalent research experience.
  • 1+ years in a research or research-engineering role.
  • Hands-on ML model training and evaluation end-to-end.
  • Strong familiarity with large language models and evaluation techniques.
  • Proficient in Python and Git; comfortable with scripting and notebooks.
  • Understanding of reinforcement learning concepts preferred.
  • Experience in AI training, model evaluation, or benchmark authoring preferred.
  • Detail-oriented, creative task design, strong written communication, independent work.
  • Availability for ~35 hours per week.

Responsibilities

  • Design tasks from real ML research ideas into defined multi-step tasks.
  • Run experiments: implement changes, train, analyze results.
  • Explore RL ideas around reward functions and training behavior.
  • Evaluate frontier models and identify gaps.
  • Collaborate with researchers to keep tasks rigorous and fair.

Skills

LLM knowledge
Experimentation experience
Reinforcement learning basics
Python & Git
Communication skills
Independent worker
Detail-oriented

Education

MSc or PhD in ML/CS

Tools

Python
Git

Job description

Cincinnatus LLC is seeking experienced ML practitioners to author complex, multi-step ML tasks for frontier models. You will design experiments, implement changes, run trainings, and analyze results to determine where models fall short.

This full-time, fully remote role within the United States requires a MSc/PhD or equivalent, 1+ years research experience, and strong Python/Git skills, with a focus on RL and LLM evaluation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Benchmark Research Scientist (Remote, Part-Time)
GenAI Benchmark Research Scientist (Remote, Part-Time)

Obsidian • San Francisco (CA)

Remote
USD 120,000 - 160,000
GenAI Benchmark Research Scientist (Remote, 35h/wk)
GenAI Benchmark Research Scientist (Remote, 35h/wk)

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 160,000
Remote Data Scientist – GenAI Benchmark & Task Design
Remote Data Scientist – GenAI Benchmark & Task Design

Obsidian • New York (NY)

On-site
USD 90,000 - 150,000
GenAI Evaluation Scientist (Remote, 35h/wk)
GenAI Evaluation Scientist (Remote, 35h/wk)

Mercor • New York (NY)

Remote
USD 90,000 - 120,000
Remote Quantitative Analyst for GenAI Benchmarking
Remote Quantitative Analyst for GenAI Benchmarking

Mercor • New York (NY)

Remote
USD 100,000 - 180,000
Remote Quantitative Analyst for AI Benchmarking
Remote Quantitative Analyst for AI Benchmarking

Mercor • New York (NY)

Remote
USD 90,000 - 130,000
GenAI Vulnerability Researcher (Remote, 35h/w)
GenAI Vulnerability Researcher (Remote, 35h/w)

Obsidian • San Francisco (CA)

Remote
USD 130,000 - 160,000
Remote GenAI Benchmark Architect — Data Science
Remote GenAI Benchmark Architect — Data Science

Mercor • New York (NY)

On-site
USD 120,000 - 170,000
GenAI Benchmark Research Scientist - Remote (35h/wk)
GenAI Benchmark Research Scientist - Remote (35h/wk)

Mercor • San Francisco (CA)

Remote
USD 120,000 - 180,000
Senior AI Engineering Specialist — GenAI Benchmarks
Senior AI Engineering Specialist — GenAI Benchmarks

Mercor • San Francisco (CA)

Hybrid
USD 180,000 - 240,000