GenAI Evaluation Scientist: LLM Benchmarks & Failures

Scale AI

New York (NY)

On-site

USD 140,000 - 210,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health coverage
PTO and flexible schedules
Learning & development stipend
Career growth opportunities

Job summary

Scale AI in New York is seeking Research Scientists and Research Engineers to advance evaluation for state-of-the-art LLMs and multimodal models. You will design benchmarks, diagnose failure modes, and publish findings from rigorous experiments.

You will collaborate with researchers and engineers across teams, translate failure analysis into actionable data and training guidance, and contribute to the next generation of scalable AI systems at Scale AI.

Qualifications

  • Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning.
  • Published research at major conferences or journals.
  • Excellent written and verbal communication skills.
  • Ph.D. or Master’s in Computer Science, ML, AI, or related field.
  • Experience in customer-facing roles is a plus.
  • Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning.

Responsibilities

  • Design and build benchmarks and evaluation methods for LLMs and multimodal models.
  • Analyze model behavior to diagnose failure modes and RCA in frontier models.
  • Collaborate with researchers and engineers to define evaluation-driven AI development.
  • Publish research findings at top-tier AI conferences.
  • Translate failure analysis into training interventions and data improvements.

Skills

RLHF
Preference modeling
Instruction tuning
LLM evaluation
Benchmark development
Written and verbal communication

Education

Ph.D. or MS in CS/ML/AI

Job description

Scale AI in New York is seeking Research Scientists and Research Engineers to advance evaluation for state-of-the-art LLMs and multimodal models. You will design benchmarks, diagnose failure modes, and publish findings from rigorous experiments.

You will collaborate with researchers and engineers across teams, translate failure analysis into actionable data and training guidance, and contribute to the next generation of scalable AI systems at Scale AI.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Evaluation Scientist: Benchmarks & Failure Insights
LLM Evaluation Scientist: Benchmarks & Failure Insights

AI Chopping Block • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Dental coverage
Vision coverage
+4
LLM Evaluation Scientist — Benchmarks & Failure Analysis
LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+3
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Base salary + equity
Health, dental, vision coverage
Retirement benefits
+3
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Retirement benefits
Learning stipend
+2
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics

Scale AI • United States

On-site
USD 181,000 - 226,000
GenAI Evaluation Scientist: Benchmarks & Failures
GenAI Evaluation Scientist: Benchmarks & Failures

Scale AI, Inc. • San Francisco (CA)

On-site
USD 181,000 - 226,000
Health benefits
Retirement benefits
Learning stipend
+2
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
AI Evaluation Engineer for LLMs & Multimodal Systems
AI Evaluation Engineer for LLMs & Multimodal Systems

ByteDance • San Jose (CA)

On-site
USD 120,000 - 190,000
Medical Insurance
Dental Insurance
Vision Insurance
+9