GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics

Scale AI, Inc.

New York, Northern (NY, KY)

Hybrid

USD 181,000 - 226,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health, dental and vision coverage
Retirement benefits
Learning and development stipend
Generous PTO
Commuter stipend

Job summary

Scale AI, Inc. invites Research Scientists and Research Engineers to join its GenAI Research Organization in New York. You will work on evaluation for LLMs and agents, building benchmarks for text and multimodal data, and applying post-training methods to improve model performance.

You will collaborate with leading labs, publish findings, and influence next-generation AI systems. The role offers base salary, equity, extensive benefits, PTO, and a commuter stipend with a salary range provided in

Qualifications

  • Ph.D. or Master’s degree in CS, ML, AI or related field.
  • Deep understanding of deep learning, RL, and large-scale model fine-tuning.
  • Experience with post-training techniques (RLHF, preference modeling, instruction tuning) and LLM evaluation or benchmarks.

Responsibilities

  • Analyze model behavior to identify and diagnose failure modes in frontier LLMs and Agents (RCA).
  • Design benchmarks and evaluation methods for text and multimodal capabilities.
  • Apply post-training techniques (SFT, RLHF, reward modeling) to connect failures to data/training interventions.
  • Publish research findings in top-tier AI conferences.

Skills

Deep learning
Reinforcement learning
LLM fine-tuning

Education

Ph.D. or Master’s in CS/ML/AI

Job description

Scale AI, Inc. invites Research Scientists and Research Engineers to join its GenAI Research Organization in New York. You will work on evaluation for LLMs and agents, building benchmarks for text and multimodal data, and applying post-training methods to improve model performance.

You will collaborate with leading labs, publish findings, and influence next-generation AI systems. The role offers base salary, equity, extensive benefits, PTO, and a commuter stipend with a salary range provided in

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Evaluation Scientist: LLM Benchmarks & Failures
GenAI Evaluation Scientist: LLM Benchmarks & Failures

Scale AI • New York (NY)

On-site
USD 140,000 - 210,000
Health coverage
PTO and flexible schedules
Learning & development stipend
+1
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Base salary + equity
Health, dental, vision coverage
Retirement benefits
+3
LLM Evaluation Scientist: Benchmarks & Failure Insights
LLM Evaluation Scientist: Benchmarks & Failure Insights

AI Chopping Block • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Dental coverage
Vision coverage
+4
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Scale AI • San Francisco (CA)

On-site
USD 166,000 - 207,000
Health coverage
Equity
Retirement benefits
+3
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Retirement benefits
Learning stipend
+2
LLM Evaluation Scientist — Benchmarks & Failure Analysis
LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+3
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics

Scale AI • United States

On-site
USD 181,000 - 226,000
GenAI Evaluation Scientist: Benchmarks & Failures
GenAI Evaluation Scientist: Benchmarks & Failures

Scale AI, Inc. • San Francisco (CA)

On-site
USD 181,000 - 226,000
Health benefits
Retirement benefits
Learning stipend
+2
GenAI Research Scientist — LLM Post-Training
GenAI Research Scientist — LLM Post-Training

Scale AI • Seattle (WA)

On-site
USD 166,000 - 207,000
Health coverage
Dental coverage
Vision coverage
+3