GenAI Evaluation Scientist: Benchmarks & Failures

Scale AI, Inc.

San Francisco (CA)

On-site

USD 181,000 - 226,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health benefits
Retirement benefits
Learning stipend
Generous PTO
Commuter stipend

Job summary

Scale AI, Inc. seeks Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation.

The role focuses on building benchmarks and diagnosing model failure modes in text and multimodal modalities within the GenAI Research Organization. You will develop rigorous evaluations, collaborate with researchers and engineers, and translate failure analysis into input for next-generation generative AI models.

Qualifications

  • PhD or master's in CS, ML, AI, or related field.
  • Deep understanding of deep learning and reinforcement learning.
  • Experience with RLHF, preference modeling, or instruction tuning, and LLM evaluation or benchmark development.
  • Published research at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR) and/or journals.
  • Excellent written and verbal communication skills.
  • Previous experience in a customer-facing role.

Responsibilities

  • Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents.
  • Design and build benchmarks and evaluation methods for text and multimodal modalities.
  • Apply post-training techniques (SFT, RLHF, reward modeling) to connect failures to data and training interventions.
  • Publish research findings in top-tier AI conferences.

Skills

Deep learning
Reinforcement learning
LLM evaluation

Education

PhD or MS in CS/ML/AI

Job description

Scale AI, Inc. seeks Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation.

The role focuses on building benchmarks and diagnosing model failure modes in text and multimodal modalities within the GenAI Research Organization. You will develop rigorous evaluations, collaborate with researchers and engineers, and translate failure analysis into input for next-generation generative AI models.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Evaluations Lead: Benchmarking & Failure Diagnostics
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics

Scale AI • United States

On-site
USD 181,000 - 226,000
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Base salary + equity
Health, dental, vision coverage
Retirement benefits
+3
LLM Evaluation Scientist — Benchmarks & Failure Analysis
LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+3
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Retirement benefits
Learning stipend
+2
LLM Evaluation Scientist: Benchmarks & Failure Insights
LLM Evaluation Scientist: Benchmarks & Failure Insights

AI Chopping Block • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Dental coverage
Vision coverage
+4
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
Generative AI Research Scientist: LLM Post-Training
Generative AI Research Scientist: LLM Post-Training

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
Machine Learning Research Scientist, Evaluations Seattle, WA Apply →
Machine Learning Research Scientist, Evaluations Seattle, WA Apply →

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
Machine Learning Research Scientist, Evaluations
Machine Learning Research Scientist, Evaluations

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2