GenAI Evaluations Lead: Benchmarking & Failure Diagnostics

Scale AI

United States

On-site

USD 181,000 - 226,000

Full time

12 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Scale AI is seeking Research Scientists and Research Engineers with expertise in LLM post-training and evaluation to join the GenAI Research Organization. The role focuses on building benchmarks, diagnosing failure modes, and evaluating models across text and multimodal modalities.

You will develop rigorous evaluation methods, collaborate with researchers and engineers, and translate failure analysis into actionable input for next-gen generative models.

Qualifications

  • Develop rigorous evaluations and diagnostic methods for frontier models.
  • Design benchmarks and evaluation methods for text and multimodal modalities.
  • Apply post-training techniques (SFT, RLHF, reward modeling) to address model failures.

Responsibilities

  • Analyze model behavior to identify failure modes in LLMs and Agents.
  • Design and build benchmarks and evaluation methods for text and multimodal tasks.
  • Apply post-training techniques to connect failures to data and interventions.
  • Publish research findings in top-tier AI conferences.

Skills

LLM evaluation
Benchmark design
Post-training methods
Communication skills

Education

Ph.D. or Master's in CS/ML/AI

Job description

Scale AI is seeking Research Scientists and Research Engineers with expertise in LLM post-training and evaluation to join the GenAI Research Organization. The role focuses on building benchmarks, diagnosing failure modes, and evaluating models across text and multimodal modalities.

You will develop rigorous evaluation methods, collaborate with researchers and engineers, and translate failure analysis into actionable input for next-gen generative models.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Evaluation Scientist: Benchmarks & Failures
GenAI Evaluation Scientist: Benchmarks & Failures

Scale AI, Inc. • San Francisco (CA)

On-site
USD 181,000 - 226,000
Health benefits
Retirement benefits
Learning stipend
+2
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Base salary + equity
Health, dental, vision coverage
Retirement benefits
+3
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Retirement benefits
Learning stipend
+2
LLM Evaluation Scientist — Benchmarks & Failure Analysis
LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+3
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Scale AI • San Francisco (CA)

On-site
USD 166,000 - 207,000
Health coverage
Equity
Retirement benefits
+3
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
LLM Evaluation Scientist: Benchmarks & Failure Insights
LLM Evaluation Scientist: Benchmarks & Failure Insights

AI Chopping Block • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Dental coverage
Vision coverage
+4
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
GenAI Research Scientist — LLM Post-Training
GenAI Research Scientist — LLM Post-Training

Scale AI • Seattle (WA)

On-site
USD 166,000 - 207,000
Health coverage
Dental coverage
Vision coverage
+3
Generative AI Research Scientist: LLM Post-Training
Generative AI Research Scientist: LLM Post-Training

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2