Senior AI Model Evaluation Scientist (LLM Benchmarks)

Cohere

Seattle (WA)

On-site

USD 180,000 - 385,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Lunch stipend
Health and dental benefits
RRSP matching / 401K / Pension
Parental leave
Enrichment benefits
Education stipend
Vacation days
Home office stipend

Job summary

Cohere is seeking a role focused on evaluating next-generation AI models, developing evaluation benchmarks and infrastructure to measure LLM progress. You will build ambitious benchmarks, work with cross-functional teams to translate model feedback into trustworthy evaluations, and advance state-of-the-art methods.

This remote-friendly role offers health and dental benefits, a lunch stipend, parental leave and education stipends, plus opportunities to work across our global offices and tools

Qualifications

  • Experience building evaluation benchmarks for ML models.
  • Ability to review complex data and model outputs.
  • Strong software engineering fundamentals.

Responsibilities

  • Create ambitious evaluation benchmarks for models.
  • Collaborate with cross-functional teams to translate feedback into evaluations.
  • Research LLM evaluation methods and improve efficiency.
  • Build scalable tools for model performance analysis.

Skills

Prototype development
Data review
Rigorous measurement
Software engineering

Job description

Cohere is seeking a role focused on evaluating next-generation AI models, developing evaluation benchmarks and infrastructure to measure LLM progress. You will build ambitious benchmarks, work with cross-functional teams to translate model feedback into trustworthy evaluations, and advance state-of-the-art methods.

This remote-friendly role offers health and dental benefits, a lunch stipend, parental leave and education stipends, plus opportunities to work across our global offices and tools

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Evaluation Scientist
Senior LLM Evaluation Scientist

cohere • New York (NY)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health benefits
RRSP matching
+5
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Scale AI • San Francisco (CA)

On-site
USD 166,000 - 207,000
Health coverage
Equity
Retirement benefits
+3
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Base salary + equity
Health, dental, vision coverage
Retirement benefits
+3
LLM Evaluation Scientist — Benchmarks & Failure Analysis
LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+3
Remote LLM Evaluation Scientist: Benchmarking Models
Remote LLM Evaluation Scientist: Benchmarking Models

Anyone AI • United States

Remote
MXN 2,626,000 - 3,678,000
Senior Research Scientist, Model Evaluation
Senior Research Scientist, Model Evaluation

SupportFinity™ • New York (NY)

On-site
USD 140,000 - 200,000
Open and inclusive culture
Work on cutting-edge AI research
Weekly lunches and snacks
+6
Senior Research Scientist, Model Evaluation (Remote‑Flexible)
Senior Research Scientist, Model Evaluation (Remote‑Flexible)

SupportFinity™ • New York (NY)

On-site
USD 140,000 - 200,000
Open and inclusive culture
Work on cutting-edge AI research
Weekly lunches and snacks
+6
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
Senior AI Engineer — LLM Evaluation & Production Systems
Senior AI Engineer — LLM Evaluation & Production Systems

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000