GenAI Evaluation Scientist - LLM Benchmarks & Failures

Scale

San Francisco, Seattle, New York (CA, WA, NY)

On-site

USD 166,000 - 207,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental coverage
Vision coverage
Retirement benefits
PTO

Job summary

Scale is seeking a Machine Learning Research Scientist, Evaluations to join the GenAI Research Organization in San Francisco. You will develop rigorous evaluations, diagnose failure modes in frontier LLMs and agents, and design benchmarks for text and multimodal modalities.

Collaboration with researchers and engineers will shape evaluation-driven AI development. The role emphasizes post-training techniques like SFT and RLHF, with opportunities to publish findings at top conferences and influence

Qualifications

  • Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field.
  • Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning.
  • Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development.
  • Excellent written and verbal communication skills.
  • Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals.
  • Previous experience in a customer facing role.

Responsibilities

  • Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents.
  • Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities.
  • Apply post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to the data and training interventions that address them.
  • Publish research findings in top-tier AI conferences.

Skills

LLM evaluation
Benchmark design
Research publication
Communication skills

Education

PhD or Master's in CS/ML

Tools

PyTorch
TensorFlow

Job description

Scale is seeking a Machine Learning Research Scientist, Evaluations to join the GenAI Research Organization in San Francisco. You will develop rigorous evaluations, diagnose failure modes in frontier LLMs and agents, and design benchmarks for text and multimodal modalities.

Collaboration with researchers and engineers will shape evaluation-driven AI development. The role emphasizes post-training techniques like SFT and RLHF, with opportunities to publish findings at top conferences and influence

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Evaluation Scientist: Benchmarks & Failures
GenAI Evaluation Scientist: Benchmarks & Failures

Scale AI, Inc. • San Francisco (CA)

On-site
USD 181,000 - 226,000
Health benefits
Retirement benefits
Learning stipend
+2
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Scale AI • San Francisco (CA)

On-site
USD 166,000 - 207,000
Health coverage
Equity
Retirement benefits
+3
LLM Evaluation Scientist — Benchmarks & Failure Analysis
LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+3
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Base salary + equity
Health, dental, vision coverage
Retirement benefits
+3
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics

Scale AI • United States

On-site
USD 181,000 - 226,000
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
LLM Evaluation Scientist: Benchmarks & Failure Insights
LLM Evaluation Scientist: Benchmarks & Failure Insights

AI Chopping Block • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Dental coverage
Vision coverage
+4
GenAI Research Engineer — LLM & Model Evaluation
GenAI Research Engineer — LLM & Model Evaluation

Google Inc. • Cambridge (MA), Northern (KY)

Hybrid
USD 174,000 - 252,000
Generative AI Research Scientist: LLM Post-Training
Generative AI Research Scientist: LLM Post-Training

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2