LLM Evaluation Scientist: Benchmarks & Failure Insights

AI Chopping Block

New York, Northern (NY, KY)

Hybrid

USD 181,000 - 226,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health coverage
Dental coverage
Vision coverage
Retirement benefits
Learning stipend
Generous PTO
Commuter stipend

Job summary

Scale is seeking Research Scientists and Research Engineers in New York to advance evaluation-driven AI development for GenAI models. You will analyze model behavior, design benchmarks, and apply post-training techniques to improve performance across text and multimodal modalities.

Join a team that collaborates with world-class labs to translate failure analysis into actionable feedback and publish findings at top conferences. Compensation includes base salary, equity, and comprehensive benefits.

Qualifications

  • PhD or Master’s degree in CS/ML or related field.
  • Deep understanding of deep learning, RL, and fine-tuning.
  • Experience with RLHF, preference modeling, or instruction tuning.

Responsibilities

  • Analyze model behavior to diagnose failure modes in frontier LLMs and Agents.
  • Design benchmarks and evaluation methods for text and multimodal modalities.
  • Apply post-training techniques to link failures to data and training interventions.
  • Publish research findings in top AI conferences.

Skills

LLM evaluation
SFT
RLHF
Reward modeling

Education

PhD or MS

Job description

Scale is seeking Research Scientists and Research Engineers in New York to advance evaluation-driven AI development for GenAI models. You will analyze model behavior, design benchmarks, and apply post-training techniques to improve performance across text and multimodal modalities.

Join a team that collaborates with world-class labs to translate failure analysis into actionable feedback and publish findings at top conferences. Compensation includes base salary, equity, and comprehensive benefits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Evaluation Scientist — Benchmarks & Failure Analysis
LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+3
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Scale AI • San Francisco (CA)

On-site
USD 166,000 - 207,000
Health coverage
Equity
Retirement benefits
+3
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Base salary + equity
Health, dental, vision coverage
Retirement benefits
+3
GenAI Evaluation Scientist - LLM Benchmarks & Failures
GenAI Evaluation Scientist - LLM Benchmarks & Failures

Scale • San Francisco (CA), Seattle (WA), New York (NY)

On-site
USD 166,000 - 207,000
Health insurance
Dental coverage
Vision coverage
+2
GenAI Evaluation Scientist: Benchmarks & Failures
GenAI Evaluation Scientist: Benchmarks & Failures

Scale AI, Inc. • San Francisco (CA)

On-site
USD 181,000 - 226,000
Health benefits
Retirement benefits
Learning stipend
+2
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics

Scale AI • United States

On-site
USD 181,000 - 226,000
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
Senior AI Model Evaluation Scientist (LLM Benchmarks)
Senior AI Model Evaluation Scientist (LLM Benchmarks)

Cohere • Seattle (WA)

On-site
USD 180,000 - 385,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K / Pension
+5
Machine Learning Research Scientist, Evaluations Seattle, WA Apply →
Machine Learning Research Scientist, Evaluations Seattle, WA Apply →

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2