GenAI Evaluation Scientist — Benchmark & Diagnose LLMs

Scale AI, Inc.

New York (NY)

On-site

USD 181,000 - 226,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Base salary + equity
Health, dental, vision coverage
Retirement benefits
Learning & development stipend
Generous PTO
Commuter stipend

Job summary

Scale AI, Inc. is recruiting for Research Scientists and Research Engineers specializing in LLM post-training evaluation and benchmark development.

This role focuses on rigorously evaluating frontier models, diagnosing failure modes, and building robust benchmarks for text and multimodal modalities. You will collaborate with researchers and engineers to define best practices in evaluation-driven AI development and translate failure analyses into strategic input for the next generation of

Qualifications

  • Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field.
  • Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning.
  • Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development.
  • Excellent written and verbal communication skills.
  • Published research in machine learning at major conferences and/or journals.
  • Previous experience in a customer facing role.

Responsibilities

  • Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents.
  • Design and build benchmarks and evaluation methods that measure LLM capabilities in text and multimodal modalities.
  • Apply post-training techniques to connect observed failures to data and training interventions.
  • Publish research findings in top-tier AI conferences.

Skills

LLM evaluation
SFT RLHF
benchmark development
communication skills
published research
customer facing experience

Education

Ph.D. or Master's in CS/ML/AI

Job description

Scale AI, Inc. is recruiting for Research Scientists and Research Engineers specializing in LLM post-training evaluation and benchmark development.

This role focuses on rigorously evaluating frontier models, diagnosing failure modes, and building robust benchmarks for text and multimodal modalities. You will collaborate with researchers and engineers to define best practices in evaluation-driven AI development and translate failure analyses into strategic input for the next generation of

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Evaluation Scientist: LLM Benchmarks & Failures
GenAI Evaluation Scientist: LLM Benchmarks & Failures

Scale AI • New York (NY)

On-site
USD 140,000 - 210,000
Health coverage
PTO and flexible schedules
Learning & development stipend
+1
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Retirement benefits
Learning stipend
+2
LLM Evaluation Scientist — Benchmarks & Failure Analysis
LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+3
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics

Scale AI • United States

On-site
USD 181,000 - 226,000
GenAI Evaluation Scientist: Benchmarks & Failures
GenAI Evaluation Scientist: Benchmarks & Failures

Scale AI, Inc. • San Francisco (CA)

On-site
USD 181,000 - 226,000
Health benefits
Retirement benefits
Learning stipend
+2
LLM Evaluation Scientist: Benchmarks & Failure Insights
LLM Evaluation Scientist: Benchmarks & Failure Insights

AI Chopping Block • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Dental coverage
Vision coverage
+4
Generative AI Research Scientist: LLM Post-Training
Generative AI Research Scientist: LLM Post-Training

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000