GenAI Evaluation Scientist (Remote, 35h/wk)

Mercor

New York (NY)

Remote

USD 90,000 - 120,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Cincinnatus LLC is seeking experienced machine learning practitioners to act as ground-truth experts for model evaluation and experimentation on a leading AI lab's GenAI team. You will author complex, multi-step ML tasks and verify where frontier models fall short.

This full-time W-2 role is fully remote in the United States, about 35 hours per week, with the opportunity to be placed at a top AI lab as part of the extended workforce.

Qualifications

  • MSc or PhD in ML, CS, or related STEM field, or equivalent research experience.
  • 1+ years in a research or research-engineering role.
  • Hands-on experience training and evaluating ML models end-to-end.
  • Strong familiarity with large language models and evaluation techniques.
  • Proficiency in Python and Git; comfortable with scripting and notebooks.
  • Basic understanding of reinforcement learning is preferred.
  • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.
  • Attention to detail, creativity, and ability to work independently in ambiguous problems.
  • Available to commit ~35 hours per week.

Responsibilities

  • Design tasks: Turn real ML research ideas into well-defined, multi-step tasks.
  • Run experiments: Implement changes, run training experiments, and analyze results.
  • Explore RL ideas: Build tasks around reinforcement-learning basics.
  • Evaluate models: See how frontier models handle tasks and note shortcomings.
  • Work as a team: Compare notes with researchers to ensure consistency and rigor.

Skills

Python
Git
LLMs
Research experience

Education

MSc or PhD in ML/CS

Job description

Cincinnatus LLC is seeking experienced machine learning practitioners to act as ground-truth experts for model evaluation and experimentation on a leading AI lab's GenAI team. You will author complex, multi-step ML tasks and verify where frontier models fall short.

This full-time W-2 role is fully remote in the United States, about 35 hours per week, with the opportunity to be placed at a top AI lab as part of the extended workforce.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Benchmark Research Scientist (Remote, Part-Time)
GenAI Benchmark Research Scientist (Remote, Part-Time)

Obsidian • San Francisco (CA)

Remote
USD 120,000 - 160,000
GenAI Benchmark Research Scientist (Remote, 35h/wk)
GenAI Benchmark Research Scientist (Remote, 35h/wk)

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 160,000
Remote Data Scientist – GenAI Benchmark & Task Design
Remote Data Scientist – GenAI Benchmark & Task Design

Obsidian • New York (NY)

On-site
USD 90,000 - 150,000
GenAI Vulnerability Researcher (Remote, 35h/w)
GenAI Vulnerability Researcher (Remote, 35h/w)

Obsidian • San Francisco (CA)

Remote
USD 130,000 - 160,000
GenAI Benchmark Research Scientist — Remote, Part-Time
GenAI Benchmark Research Scientist — Remote, Part-Time

Obsidian • New York (NY)

Remote
USD 120,000 - 150,000
Remote GenAI Benchmark Architect — Data Science
Remote GenAI Benchmark Architect — Data Science

Mercor • New York (NY)

On-site
USD 120,000 - 170,000
Finance AI Model Evaluation Specialist
Finance AI Model Evaluation Specialist

Mercor • United States

On-site
USD 180,000 - 280,000
GenAI Marketing Strategist & Model Evaluator
GenAI Marketing Strategist & Model Evaluator

Mercor • United States

On-site
USD 120,000 - 180,000
Finance AI Evaluation Expert for GenAI Models
Finance AI Evaluation Expert for GenAI Models

Mercor • New York (NY)

On-site
USD 150,000 - 210,000
GenAI Insurance SME — Risk & Underwriting Evaluation
GenAI Insurance SME — Risk & Underwriting Evaluation

Obsidian • New York (NY)

On-site
USD 140,000 - 220,000