ML Evaluation Scientist - LLM Benchmarks

Scale AI

Seattle (WA)

On-site

USD 181,000 - 226,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health coverage
Retirement benefits
L&D stipend
Generous PTO
Commuter stipend

Job summary

Scale AI is hiring for roles focused on evaluating and benchmarking frontier LLMs and Agents within the GenAI Research Organization. You will develop robust evaluations and RCA-driven insights, collaborating with researchers to shape evaluation-driven AI development and translate failure analyses into strategic input for next-gen models.

The role emphasizes post-training techniques like SFT and RLHF, publication of findings at major AI conferences, and building benchmarks for text and multimodal

Qualifications

  • Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field.
  • Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning.
  • Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development.
  • Excellent written and verbal communication skills.
  • Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals.
  • Previous experience in a customer facing role.

Responsibilities

  • Analyze model behavior to identify failure modes in frontier LLMs and Agents.
  • Design benchmarks and evaluation methods for text and multimodal modalities.
  • Apply post-training techniques to connect observed failures to data and training interventions.
  • Publish research findings in top-tier AI conferences.

Skills

LLM evaluation
RLHF
SFT
Reward modeling
Diagnostics

Education

PhD or MSc in CS/ML

Tools

Python
PyTorch
Benchmark tools

Job description

Scale AI is hiring for roles focused on evaluating and benchmarking frontier LLMs and Agents within the GenAI Research Organization. You will develop robust evaluations and RCA-driven insights, collaborating with researchers to shape evaluation-driven AI development and translate failure analyses into strategic input for next-gen models.

The role emphasizes post-training techniques like SFT and RLHF, publication of findings at major AI conferences, and building benchmarks for text and multimodal

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Evaluation Scientist — Benchmark & Diagnose LLMs
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Base salary + equity
Health, dental, vision coverage
Retirement benefits
+3
LLM Evaluation Scientist — Benchmarks & Failure Analysis
LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+3
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Retirement benefits
Learning stipend
+2
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
LLM Evaluation Scientist: Benchmarks & Failure Insights
LLM Evaluation Scientist: Benchmarks & Failure Insights

AI Chopping Block • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Dental coverage
Vision coverage
+4
Remote LLM Evaluation Scientist: Benchmarking Models
Remote LLM Evaluation Scientist: Benchmarking Models

Anyone AI • United States

Remote
MXN 2,626,000 - 3,678,000
GenAI Evaluation Scientist: Benchmarks & Failures
GenAI Evaluation Scientist: Benchmarks & Failures

Scale AI, Inc. • San Francisco (CA)

On-site
USD 181,000 - 226,000
Health benefits
Retirement benefits
Learning stipend
+2
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
AI Evaluation Engineer for LLMs & Multimodal Systems
AI Evaluation Engineer for LLMs & Multimodal Systems

ByteDance • San Jose (CA)

On-site
USD 120,000 - 190,000
Medical Insurance
Dental Insurance
Vision Insurance
+9
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics

Scale AI • United States

On-site
USD 181,000 - 226,000