Senior Model Evaluation Engineer – LLM Benchmarking

cohere

New York (NY)

Hybrid

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

A weekly lunch stipend of $75/£75 or?e
Full health and dental benefits
RRSP matching, 401K, Pension Scheme
100% Parental Leave top‑up
Annual enrichment benefits and cowork/

Job summary

Cohere is seeking a Senior Research Engineer, Model Evaluation, to create next‑generation evaluation methods and scalable infrastructure. You will develop benchmarks, datasets, and environments to measure frontier model capabilities, and you will push the state‑of‑the‑art in LLM evaluation methods while building tools used across technical staff and leadership.

You will collaborate with leading researchers and engineers, focusing on improving evaluation efficiency and scalable dataset

Qualifications

  • You enjoy pushing the limits of LLM evaluation and have built high‑quality evaluation resources (datasets, simulators, environments).
  • You have a track record of developing new methods or data to evaluate LLMs, e.g. publications or benchmarks.
  • You have deep experience building with and around LLMs and you’ve built tools to analyze performance.

Responsibilities

  • Develop evaluation benchmarks, datasets, and environments for measuring cutting‑edge model capabilities.
  • Conduct research to advance LLM evaluation methods, including training LLM judges and efficient evaluation.
  • Build scalable tools to investigate and understand evaluation results for the org and leadership.
  • Collaborate with top researchers and engineers to push state‑of‑the‑art in evaluation.

Skills

LLM evaluation
Research engineering
Software engineering
Publications / benchmarks

Tools

Python
PyTorch / TensorFlow
Datasets & evaluation tooling

Job description

Cohere is seeking a Senior Research Engineer, Model Evaluation, to create next‑generation evaluation methods and scalable infrastructure. You will develop benchmarks, datasets, and environments to measure frontier model capabilities, and you will push the state‑of‑the‑art in LLM evaluation methods while building tools used across technical staff and leadership.

You will collaborate with leading researchers and engineers, focusing on improving evaluation efficiency and scalable dataset

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Evaluation Scientist
Senior LLM Evaluation Scientist

cohere • New York (NY)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health benefits
RRSP matching
+5
Senior Research Scientist, Model Evaluation (Remote‑Flexible)
Senior Research Scientist, Model Evaluation (Remote‑Flexible)

SupportFinity™ • New York (NY)

On-site
USD 140,000 - 200,000
Open and inclusive culture
Work on cutting-edge AI research
Weekly lunches and snacks
+6
Senior LLM Evaluation Infrastructure Engineer
Senior LLM Evaluation Infrastructure Engineer

Inception • San Francisco (CA)

On-site
USD 200,000 - 350,000
Health insurance
Equity
Flexible vacation
+2
LLM Evaluations Engineer — Benchmark Leaderboards
LLM Evaluations Engineer — Benchmark Leaderboards

Vals AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health/dental insurance coverage
Relocation support
Lunch and dinner provided
+2
LLM Benchmark Engineer — Lead Leaderboard Insights
LLM Benchmark Engineer — Lead Leaderboard Insights

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 190,000
Relocation and transportation support
Health/dental insurance coverage
Lunch and dinner provided
+2
Remote LLM Evaluation Scientist: Benchmarking Models
Remote LLM Evaluation Scientist: Benchmarking Models

Anyone AI • United States

Remote
MXN 2,626,000 - 3,678,000
Senior Research Engineer, Model Evaluation
Senior Research Engineer, Model Evaluation

cohere • New York (NY)

Hybrid
USD 150,000 - 210,000
A weekly lunch stipend of $75/£75 or?e
Full health and dental benefits
RRSP matching, 401K, Pension Scheme
+2
Member of Technical Staff - Research
Member of Technical Staff - Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Relocation support
Health insurance
Meals provided
+2
ML Evaluation Engineer: Benchmark & Model Quality
ML Evaluation Engineer: Benchmark & Model Quality

Reducto • San Francisco (CA)

On-site
USD 100,000 - 130,000
Unlimited PTO
Daily free lunch
Reimbursed transportation
+3
Senior LLM Infra Engineer — HPC, Benchmarking & DevEx
Senior LLM Infra Engineer — HPC, Benchmarking & DevEx

Baseten • United States

Remote
USD 180,000 - 300,000