GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Neura Market

San Francisco, Northern (CA, KY)

Hybrid

USD 181,000 - 226,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health coverage
Retirement benefits
Learning stipend
Generous PTO
Commuter stipend

Job summary

Scale is seeking Research Scientists and Research Engineers in San Francisco to build rigorous evaluations and benchmarks for LLMs and multimodal models. You will diagnose failure modes, design evaluation methods, and apply post-training techniques to guide data interventions.

You will publish findings with top researchers and partner with leading labs to shape next-gen generative AI. The role emphasizes collaboration with researchers and engineers, cross-functional communication, and

Qualifications

  • PhD or Master's in CS/ML/AI or related field.
  • Deep understanding of deep learning, RL, and large-scale model fine-tuning.
  • Experience with RLHF, reward modeling, instruction tuning, or LLM evaluation.
  • Strong written and verbal communication skills.
  • Published research at NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, or similar.
  • Experience in customer-facing roles.

Responsibilities

  • Analyze model behavior to identify failure modes in frontier LLMs and Agents.
  • Design benchmarks and evaluation methods for text and multimodal outputs.
  • Apply post-training techniques (SFT, RLHF, reward modeling) to connect failures to data interventions.
  • Publish research findings at top AI conferences.

Skills

Deep learning
Reinforcement learning
LLM evaluation
Communication skills
Research publication experience
Customer facing experience

Education

PhD or Master's degree in CS/ML/AI

Job description

Scale is seeking Research Scientists and Research Engineers in San Francisco to build rigorous evaluations and benchmarks for LLMs and multimodal models. You will diagnose failure modes, design evaluation methods, and apply post-training techniques to guide data interventions.

You will publish findings with top researchers and partner with leading labs to shape next-gen generative AI. The role emphasizes collaboration with researchers and engineers, cross-functional communication, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Evaluation Scientist — Benchmark & Diagnose LLMs
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Base salary + equity
Health, dental, vision coverage
Retirement benefits
+3
LLM Evaluation Scientist — Benchmarks & Failure Analysis
LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+3
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics
GenAI Evaluations Lead: Benchmarking & Failure Diagnostics

Scale AI • United States

On-site
USD 181,000 - 226,000
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
GenAI Evaluation Scientist: Benchmarks & Failures
GenAI Evaluation Scientist: Benchmarks & Failures

Scale AI, Inc. • San Francisco (CA)

On-site
USD 181,000 - 226,000
Health benefits
Retirement benefits
Learning stipend
+2
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
LLM Evaluation Scientist: Benchmarks & Failure Insights
LLM Evaluation Scientist: Benchmarks & Failure Insights

AI Chopping Block • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Dental coverage
Vision coverage
+4
Generative AI Research Scientist: LLM Post-Training
Generative AI Research Scientist: LLM Post-Training

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
LLM Benchmark Lead Research Scientist
LLM Benchmark Lead Research Scientist

Vals AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Relocation support
Health insurance
Lunch and snacks provided
+2
LLM Post-Training Research Scientist
LLM Post-Training Research Scientist

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2