GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Scale AI

San Francisco (CA)

On-site

USD 166,000 - 207,000

Full time

13 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health coverage
Equity
Retirement benefits
Learning stipend
Generous PTO
Commuter stipend

Job summary

Scale AI is seeking Research Scientists and Research Engineers to advance evaluation for LLMs and multimodal models. You will analyze failures, design rigorous benchmarks, and connect findings to data and training interventions.

You’ll publish results at top AI conferences and collaborate with leading labs to influence future model development. The role emphasizes strong communication, deep learning expertise, and experience with post-training techniques like RLHF and preference modeling in

Qualifications

  • Ph.D. or Master’s degree in Computer Science, Machine Learning, AI, or a related field.
  • Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning.
  • Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development.
  • Excellent written and verbal communication skills.

Responsibilities

  • Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents.
  • Design and build benchmarks and evaluation methods that measure LLM capabilities in text and multimodal modalities.
  • Apply post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to data and training interventions.
  • Publish research findings in top-tier AI conferences.

Skills

LLM evaluation
RCA (root cause analysis)
Multimodal evaluation
Benchmarks design
Research publication
Communication skills

Education

Ph.D. or Master’s in Computer Science, ML, AI

Job description

Scale AI is seeking Research Scientists and Research Engineers to advance evaluation for LLMs and multimodal models. You will analyze failures, design rigorous benchmarks, and connect findings to data and training interventions.

You’ll publish results at top AI conferences and collaborate with leading labs to influence future model development. The role emphasizes strong communication, deep learning expertise, and experience with post-training techniques like RLHF and preference modeling in

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Evaluation Scientist — Benchmark & Diagnose LLMs
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Base salary + equity
Health, dental, vision coverage
Retirement benefits
+3
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Retirement benefits
Learning stipend
+2
LLM Evaluation Scientist — Benchmarks & Failure Analysis
LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+3
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
GenAI Evaluation Scientist: Benchmarks & Failures
GenAI Evaluation Scientist: Benchmarks & Failures

Scale AI, Inc. • San Francisco (CA)

On-site
USD 181,000 - 226,000
Health benefits
Retirement benefits
Learning stipend
+2
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics

Scale AI • United States

On-site
USD 181,000 - 226,000
LLM Evaluation Scientist: Benchmarks & Failure Insights
LLM Evaluation Scientist: Benchmarks & Failure Insights

AI Chopping Block • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Dental coverage
Vision coverage
+4
GenAI Research Scientist — LLM Post-Training
GenAI Research Scientist — LLM Post-Training

Scale AI • Seattle (WA)

On-site
USD 166,000 - 207,000
Health coverage
Dental coverage
Vision coverage
+3
Generative AI Research Scientist: LLM Post-Training
Generative AI Research Scientist: LLM Post-Training

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2