LLM Benchmarking Research Scientist

Anyone AI Inc.

Northern (KY)

Hybrid

USD 110,000 - 160,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Anyone AI Labs in a remote capacity across LatAm/US seeks a Research Scientist focused on LLM evaluations and benchmarking. You will design frontier-grade evaluation methods and build benchmarks across reasoning, coding, agents, and multi-modal capabilities.

You will collaborate with labs, defend methodologies under scrutiny, and contribute to public benchmarks and papers, targeting venues like NeurIPS Datasets & Benchmarks, ICLR, and ACL.

Qualifications

  • Research background in ML evaluation or benchmarking.
  • Expertise in LLM/frontier-model benchmarking.
  • Understanding measurement constructs and psychometrics.
  • Interest in safety and robustness of evaluation.
  • Ability to lead a team and maintain high standards.
  • Experience framing, running studies, and writing results.

Responsibilities

  • Evaluation research: convert targets into benchmark designs.
  • Benchmark development with expert-verified ground truth.
  • Recruit and calibrate a pool of experts across coding, agents, and reasoning.
  • Serve as the technical contact for labs with CEO support.
  • Turn lab requests into pilots and publish benchmarks.

Skills

ML evaluation
LLM benchmarking
Measurement science
Safety in evaluation
Team leadership
Study design

Job description

Anyone AI Labs in a remote capacity across LatAm/US seeks a Research Scientist focused on LLM evaluations and benchmarking. You will design frontier-grade evaluation methods and build benchmarks across reasoning, coding, agents, and multi-modal capabilities.

You will collaborate with labs, defend methodologies under scrutiny, and contribute to public benchmarks and papers, targeting venues like NeurIPS Datasets & Benchmarks, ICLR, and ACL.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Scientist
Research Scientist

Anyone AI Inc. • Northern (KY)

Hybrid
USD 110,000 - 160,000
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
LLM Evaluation Scientist — Benchmarks & Failure Analysis
LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+3
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Scale AI • San Francisco (CA)

On-site
USD 166,000 - 207,000
Health coverage
Equity
Retirement benefits
+3
Senior AI Model Evaluation Scientist (LLM Benchmarks)
Senior AI Model Evaluation Scientist (LLM Benchmarks)

Cohere • Seattle (WA)

On-site
USD 180,000 - 385,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K / Pension
+5
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Base salary + equity
Health, dental, vision coverage
Retirement benefits
+3
LLM Benchmark Engineer — Lead Leaderboard Insights
LLM Benchmark Engineer — Lead Leaderboard Insights

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 190,000
Relocation and transportation support
Health/dental insurance coverage
Lunch and dinner provided
+2
GenAI Evaluation Scientist - LLM Benchmarks & Failures
GenAI Evaluation Scientist - LLM Benchmarks & Failures

Scale • San Francisco (CA), Seattle (WA), New York (NY)

On-site
USD 166,000 - 207,000
Health insurance
Dental coverage
Vision coverage
+2
LLM Benchmark Architect — Research Lead
LLM Benchmark Architect — Research Lead

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 185,000
Relocation assistance
Housing stipend (within 1 mile)
Health and dental insurance
+2
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2