Senior Research Scientist, Model Evaluation (Remote‑Flexible)

SupportFinity™

New York (NY)

On-site

USD 140,000 - 200,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Open and inclusive culture
Work on cutting-edge AI research
Weekly lunches and snacks
Full health and dental benefits
Parental Leave top-up
Enrichment perks for arts and wellness
Remote-flexible, offices in multiple'
Co-working stipend
Six weeks of vacation

Job summary

Cohere is seeking a Senior Research Scientist, Model Evaluation in New York to lead the development of next-generation evaluation methods and infrastructure for measuring LLM progress. You will design ambitious benchmarks and work across teams to translate model feedback into robust, repeatable assessments.

The role emphasizes research, evaluation efficiency, and building scalable tooling to probe model performance, with a focus on aligning measurements with high-value capabilities.

Qualifications

  • Experience creating evaluation benchmarks for AI/ML models.
  • Experience translating model feedback into repeatable evaluations.
  • Research to advance state-of-the-art in LLM evaluation methods.
  • Strong software engineering skills.

Responsibilities

  • Create ambitious new evaluation benchmarks that push model capabilities.
  • Collaborate with cross-functional teams to translate feedback into trustworthy evaluations.
  • Conduct research to advance LLM evaluation methods and improve efficiency.
  • Build scalable, reusable tools for analyzing model performance.

Skills

Evaluation benchmarks
LLM training
LLM data synthesis
Evaluation tooling
Software engineering

Tools

Data synthesis pipelines
Scalable tooling

Job description

Cohere is seeking a Senior Research Scientist, Model Evaluation in New York to lead the development of next-generation evaluation methods and infrastructure for measuring LLM progress. You will design ambitious benchmarks and work across teams to translate model feedback into robust, repeatable assessments.

The role emphasizes research, evaluation efficiency, and building scalable tooling to probe model performance, with a focus on aligning measurements with high-value capabilities.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior LLM Evaluation Scientist
Senior LLM Evaluation Scientist

cohere • New York (NY)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health benefits
RRSP matching
+5
Senior AI Model Evaluation Scientist (LLM Benchmarks)
Senior AI Model Evaluation Scientist (LLM Benchmarks)

Cohere • Seattle (WA)

On-site
USD 180,000 - 385,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K / Pension
+5
Senior Research Scientist, Model Evaluation
Senior Research Scientist, Model Evaluation

SupportFinity™ • New York (NY)

On-site
USD 140,000 - 200,000
Open and inclusive culture
Work on cutting-edge AI research
Weekly lunches and snacks
+6
Senior Research Scientist, Model Evaluation
Senior Research Scientist, Model Evaluation

cohere • New York (NY)

On-site
USD 150,000 - 210,000
Lunch stipend
Health benefits
RRSP matching
+5
LLM Evaluation Scientist: Benchmarks & Failure Insights
LLM Evaluation Scientist: Benchmarks & Failure Insights

AI Chopping Block • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Dental coverage
Vision coverage
+4
Senior Research Scientist, Model Evaluation
Senior Research Scientist, Model Evaluation

Cohere • Seattle (WA)

On-site
USD 180,000 - 385,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K / Pension
+5
Research Scientist
Research Scientist

Anyone AI Inc. • Northern (KY)

Hybrid
USD 110,000 - 160,000
LLM Evaluations Engineer — Remote Benchmarking & Tools
LLM Evaluations Engineer — Remote Benchmarking & Tools

Jaide Health • United States

On-site
USD 90,000 - 130,000
Fully remote work & flexible hours
37 days/year of vacation & holidays
Health insurance allowance for you and dependents
+4
Data Analysis & Evaluation Engineer (Remote)
Data Analysis & Evaluation Engineer (Remote)

Cohere • United States

Remote
USD 140,000 - 210,000
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Scale AI • San Francisco (CA)

On-site
USD 166,000 - 207,000
Health coverage
Equity
Retirement benefits
+3