Remote LLM Evaluation Scientist: Benchmarking Models

Anyone AI

United States

Remote

MXN 2,626,510 - 3,677,114

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anyone AI Labs is seeking a Research Scientist for LLM Evaluations and Benchmarking. This remote role spans LatAm/US and requires designing robust evaluation methodologies for frontier models, building benchmarks across reasoning, coding, agents, tool use, and multi-modal tasks.

The candidate will lead expert pools, validate ground truth, and publish results in venues like NeurIPS, ICLR, and ACL, with strong English proficiency and preference for Spanish speakers.

Qualifications

  • Experience in ML evaluation or benchmarking with published benchmarks.
  • Deep knowledge of LLM/frontier-model benchmarking, including code and agentic eval.
  • Understanding of construct validity, psychometrics, rubrics, headroom.
  • Interest in safety and robustness in evaluation.
  • Ability to lead a team of experts and run full research cycles.

Responsibilities

  • Research evaluation methods and design new benchmarks.
  • Develop evaluation packages with ground truth and multi-model headroom.
  • Recruit and calibrate a pool of experts across coding, agents, STEM.
  • Serve as technical point of contact for labs with CEO support.
  • Turn lab requests into pilot studies and publish benchmarks/papers.

Skills

ML evaluation
Benchmarking
Frontier-model benchmarking
Code-model evaluation
Agentic evaluation
Measurement theory
English fluency

Job description

Anyone AI Labs is seeking a Research Scientist for LLM Evaluations and Benchmarking. This remote role spans LatAm/US and requires designing robust evaluation methodologies for frontier models, building benchmarks across reasoning, coding, agents, tool use, and multi-modal tasks.

The candidate will lead expert pools, validate ground truth, and publish results in venues like NeurIPS, ICLR, and ACL, with strong English proficiency and preference for Spanish speakers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist (Remote/US/LATAM)
Research Scientist (Remote/US/LATAM)

Anyone AI • United States

Remote
MXN 2,626,000 - 3,678,000
Remote AI Math Evaluator & LLM Benchmark Designer
Remote AI Math Evaluator & LLM Benchmark Designer

United States Digital Space LLC • United States

Remote
Fully remote
AI projects
Contract extension potential
Senior LLM Evaluation Scientist
Senior LLM Evaluation Scientist

cohere • New York (NY)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health benefits
RRSP matching
+5
Remote AI Research Scientist — Applied LLM Evaluation
Remote AI Research Scientist — Applied LLM Evaluation

AIToolboard • New York (NY)

Remote
AI Research Scientist (United States, Remote)
AI Research Scientist (United States, Remote)

Rex.zone • United States

Remote
Senior Research Scientist, Model Evaluation (Remote‑Flexible)
Senior Research Scientist, Model Evaluation (Remote‑Flexible)

SupportFinity™ • New York (NY)

On-site
USD 140,000 - 200,000
Open and inclusive culture
Work on cutting-edge AI research
Weekly lunches and snacks
+6
Remote Applied AI Research Scientist (LLM & Evaluation)
Remote Applied AI Research Scientist (LLM & Evaluation)

Rex.zone • United States

Remote
USD 80,000 - 100,000
Remote AI Benchmark Engineer & Researcher
Remote AI Benchmark Engineer & Researcher

Pathway • Palo Alto (CA)

Remote
USD 120,000 - 180,000
AI Engineer - LLM Training & Evaluation (Remote)
AI Engineer - LLM Training & Evaluation (Remote)

Prolific • Memphis (TN)

Hybrid
USD 90,000 - 130,000
Competitive pay rates
Flexible hours
Ability to work from home
LLM Evaluations Engineer — Benchmark Leaderboards
LLM Evaluations Engineer — Benchmark Leaderboards

Vals AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health/dental insurance coverage
Relocation support
Lunch and dinner provided
+2