Research Scientist (Remote/US/LATAM)

Anyone AI

United States

Remote

MXN 2,626,510 - 3,677,114

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Anyone AI Labs is seeking a Research Scientist for LLM Evaluations and Benchmarking. This remote role spans LatAm/US and requires designing robust evaluation methodologies for frontier models, building benchmarks across reasoning, coding, agents, tool use, and multi-modal tasks.

The candidate will lead expert pools, validate ground truth, and publish results in venues like NeurIPS, ICLR, and ACL, with strong English proficiency and preference for Spanish speakers.

Qualifications

  • Experience in ML evaluation or benchmarking with published benchmarks.
  • Deep knowledge of LLM/frontier-model benchmarking, including code and agentic eval.
  • Understanding of construct validity, psychometrics, rubrics, headroom.
  • Interest in safety and robustness in evaluation.
  • Ability to lead a team of experts and run full research cycles.

Responsibilities

  • Research evaluation methods and design new benchmarks.
  • Develop evaluation packages with ground truth and multi-model headroom.
  • Recruit and calibrate a pool of experts across coding, agents, STEM.
  • Serve as technical point of contact for labs with CEO support.
  • Turn lab requests into pilot studies and publish benchmarks/papers.

Skills

ML evaluation
Benchmarking
Frontier-model benchmarking
Code-model evaluation
Agentic evaluation
Measurement theory
English fluency

Job description

Research Scientist, LLM Evaluations & Benchmarking

Anyone AI Labs

Remote / LatAm / US

The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or measuring the wrong thing. You'll own that problem at Anyone AI, measuring frontier model capability. This is a research role at heart: you decide what a good evaluation is, design the benchmarks that prove it, and defend the methodology under lab scrutiny. You'll build frontier-grade evaluation packages across reasoning, coding, agents, tool use, and multi-modal — grounded in expert-verified truth, validated against multiple models, and QC'd to survive buyer-side review.

Responsibilities
  • Evaluation research. Turn eval targets into original benchmark designs. Own the hard measurement questions: construct validity, item discrimination, headroom, reliability, contamination, and capability elicitation. Push toward evals that stay informative as models improve.
  • Benchmark development. Build evaluation packages with subject‑matter experts, each with expert‑verified ground truth, multi‑model headroom results, and rigorous QC (calibration layers, severity‑weighted rubrics, deterministic verifiers).
  • Experts. Recruit, calibrate, and review a pool across coding, agentic/tool‑use, and STEM/reasoning. Be the final arbiter of correctness and frontier difficulty.
  • Lab relationships. Be a technical point of contact for labs, with CEO support. Understand what they're trying to measure and translate it into an evaluation design.
  • Delivery & dissemination. Turn lab requests into winning sample packages and own pilots end to end. Where the work generalizes, help turn it into public benchmarks and papers: we support publishing at venues like NeurIPS Datasets & Benchmarks, ICLR, and ACL.
What We're Looking For
  • Research background in ML evaluation or benchmarking (a track record of published or open benchmarks, eval/measurement research, or equivalent hands‑on work that labs have relied on).
  • Deep LLM/frontier-model benchmarking expertise, with real strength in code‑model and agentic evaluation.
  • Fluency with the measurement problem itself: construct validity, psychometrics, rubrics, pass rates, headroom, contamination, and what makes a task genuinely discriminate a model.
  • Interest in the safety side of evaluation, capability elicitation, robustness, and measuring the things that are hardest to measure honestly.
  • Proven ability to hold a team or expert pool to a rigorous standard.
  • Comfort with the full research loop: framing the question, running the study, and writing it up.
  • Fluent English; Spanish a plus.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote LLM Evaluation Scientist: Benchmarking Models
Remote LLM Evaluation Scientist: Benchmarking Models

Anyone AI • United States

Remote
MXN 2,626,000 - 3,678,000
Member of Technical Staff - Research
Member of Technical Staff - Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Relocation support
Health insurance
Meals provided
+2
AI Research scientist - Evals
AI Research scientist - Evals

Cerebro • San Francisco (CA)

On-site
USD 165,000 - 195,000
Equity
Relocation support
Health and dental insurance
+2
Head of Research
Head of Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 220,000 - 340,000
Relocation support
Health insurance
Meals provided (Lunch/Dinner)
+3
Research Scientist, Agentic Data & Benchmarking
Research Scientist, Agentic Data & Benchmarking

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 180,000 - 250,000
Member of Technical Staff - Research
Member of Technical Staff - Research

Vals AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Relocation support
Health insurance
Lunch and snacks provided
+2
AI Research Scientist (United States, Remote)
AI Research Scientist (United States, Remote)

Rex.zone • United States

Remote
Head of Research
Head of Research

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 225,000 - 275,000
Relocation and transportation support
Health and dental insurance
Lunch and dinner provided
+6
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
Member of Technical Staff - Research
Member of Technical Staff - Research

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 185,000
Relocation assistance
Housing stipend (within 1 mile)
Health and dental insurance
+2