Senior LLM Evaluation Scientist

cohere

New York (NY)

Hybrid

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Lunch stipend
Health benefits
RRSP matching
Parental leave
Learning stipend
Vacation time
Travel to offices
Home office stipend

Job summary

Cohere is seeking a Senior Research Scientist, Model Evaluation, to create ambitious evaluation benchmarks and scale evaluation infrastructure for enterprise AI. You will work with cross-functional teams to translate model feedback into trustworthy measurements and to push the frontiers of LLM evaluation.

The role emphasizes rigorous measurement, practical tooling, and collaboration across research and engineering, with remote-friendly policies and opportunities to influence model evaluation at

Qualifications

  • You enjoy rapidly building prototypes that demonstrate the boundaries of what LLMs are capable of, and you have developed resources to measure those capabilities.
  • You have spent dozens of hours reviewing complex data and LLM outputs to ensure high data quality.
  • You are obsessive about rigorously measuring AI capabilities, and also about making sure your measurements actually align with the capabilities you care about.
  • You have strong software engineering skills.

Responsibilities

  • Create ambitious new evaluation benchmarks that push the limits of what our models can accomplish.
  • Work on highly cross-functional teams to translate model feedback into trustworthy, repeatable evaluations.
  • Conduct research to advance the state-of-the-art in LLM evaluation methods, including training LLM judges; refining LLM-based data synthesis pipelines; and improving evaluation efficiency.
  • Build scalable and reusable tools for digging into model performance.

Skills

Prototype building
LLM evaluation
Data quality
Software engineering

Job description

Cohere is seeking a Senior Research Scientist, Model Evaluation, to create ambitious evaluation benchmarks and scale evaluation infrastructure for enterprise AI. You will work with cross-functional teams to translate model feedback into trustworthy measurements and to push the frontiers of LLM evaluation.

The role emphasizes rigorous measurement, practical tooling, and collaboration across research and engineering, with remote-friendly policies and opportunities to influence model evaluation at

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Model Evaluation Engineer – LLM Benchmarking
Senior Model Evaluation Engineer – LLM Benchmarking

cohere • New York (NY)

Hybrid
USD 150,000 - 210,000
A weekly lunch stipend of $75/£75 or?e
Full health and dental benefits
RRSP matching, 401K, Pension Scheme
+2
Senior Research Scientist, Model Evaluation (Remote‑Flexible)
Senior Research Scientist, Model Evaluation (Remote‑Flexible)

SupportFinity™ • New York (NY)

On-site
USD 140,000 - 200,000
Open and inclusive culture
Work on cutting-edge AI research
Weekly lunches and snacks
+6
Senior Research Scientist, Model Evaluation
Senior Research Scientist, Model Evaluation

SupportFinity™ • New York (NY)

On-site
USD 140,000 - 200,000
Open and inclusive culture
Work on cutting-edge AI research
Weekly lunches and snacks
+6
Remote LLM Evaluation Scientist: Benchmarking Models
Remote LLM Evaluation Scientist: Benchmarking Models

Anyone AI • United States

Remote
MXN 2,626,000 - 3,678,000
Senior Applied AI Scientist - LLM Evaluation
Senior Applied AI Scientist - LLM Evaluation

Dadi Inc (acquired by Ro) • New York (NY)

On-site
USD 182,000 - 220,000
Competitive equity
Benefits package
Health benefits
Senior Research Engineer, Model Evaluation
Senior Research Engineer, Model Evaluation

cohere • New York (NY)

Hybrid
USD 150,000 - 210,000
A weekly lunch stipend of $75/£75 or?e
Full health and dental benefits
RRSP matching, 401K, Pension Scheme
+2
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
Staff Data Scientist — LLM Data Analysis & Evaluation
Staff Data Scientist — LLM Data Analysis & Evaluation

Cohere • New York (NY), San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health benefits
RRSP matching
+6
Senior AI Engineer — LLM Evaluation & Production Systems
Senior AI Engineer — LLM Evaluation & Production Systems

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000
Research Scientist (Remote/US/LATAM)
Research Scientist (Remote/US/LATAM)

Anyone AI • United States

Remote
MXN 2,626,000 - 3,678,000