LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc.

Seattle (WA)

On-site

USD 181,000 - 226,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental and vision coverage
Retirement benefits
Learning and development stipend
Generous PTO
Commuter stipend
Equity-based compensation

Job summary

Scale AI, Inc. is seeking Research Scientists and Research Engineers to join the GenAI Research Organization. You will focus on evaluation of LLMs and multimodal models, building benchmarks, and diagnosing failure modes to guide improvements.

You will collaborate with researchers and engineers, apply post-training techniques like SFT and RLHF, publish results at top conferences, and help shape the next generation of generative AI models, with competitive compensation and equity.

Qualifications

  • PhD or MS in Computer Science, ML, AI, or a related field.
  • Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning.
  • Experience with RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development.
  • Excellent written and verbal communication skills.
  • Published research in machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR) and/or journals.
  • Previous experience in a customer facing role.

Responsibilities

  • Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents; RCA focused.
  • Design and build benchmarks and evaluation methods measuring LLM capabilities in text and multimodal modalities.
  • Apply post-training techniques (SFT, RLHF, reward modeling) to connect failures to data and training interventions.
  • Publish research findings in top-tier AI conferences.

Skills

LLM evaluation
Post-training techniques
Benchmark development
Communication skills
Research publications
Customer facing

Education

PhD or MS in CS/ML/AI

Tools

PyTorch
TensorFlow
Python

Job description

Scale AI, Inc. is seeking Research Scientists and Research Engineers to join the GenAI Research Organization. You will focus on evaluation of LLMs and multimodal models, building benchmarks, and diagnosing failure modes to guide improvements.

You will collaborate with researchers and engineers, apply post-training techniques like SFT and RLHF, publish results at top conferences, and help shape the next generation of generative AI models, with competitive compensation and equity.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Evaluation Scientist — Benchmark & Diagnose LLMs
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Base salary + equity
Health, dental, vision coverage
Retirement benefits
+3
LLM Evaluation Scientist: Benchmarks & Failure Insights
LLM Evaluation Scientist: Benchmarks & Failure Insights

AI Chopping Block • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Dental coverage
Vision coverage
+4
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Retirement benefits
Learning stipend
+2
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics

Scale AI • United States

On-site
USD 181,000 - 226,000
GenAI Evaluation Scientist: Benchmarks & Failures
GenAI Evaluation Scientist: Benchmarks & Failures

Scale AI, Inc. • San Francisco (CA)

On-site
USD 181,000 - 226,000
Health benefits
Retirement benefits
Learning stipend
+2
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
Generative AI Research Scientist: LLM Post-Training
Generative AI Research Scientist: LLM Post-Training

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
LLM Post-Training Research Scientist
LLM Post-Training Research Scientist

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
LLM Post-Training Research Scientist (SFT & RLHF)
LLM Post-Training Research Scientist (SFT & RLHF)

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Health, dental & vision coverage
Retirement benefits
Learning and development stipend
+2