Research Scientist

Anyone AI Inc.

Northern (KY)

Hybrid

USD 110,000 - 160,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Anyone AI Labs in a remote capacity across LatAm/US seeks a Research Scientist focused on LLM evaluations and benchmarking. You will design frontier-grade evaluation methods and build benchmarks across reasoning, coding, agents, and multi-modal capabilities.

You will collaborate with labs, defend methodologies under scrutiny, and contribute to public benchmarks and papers, targeting venues like NeurIPS Datasets & Benchmarks, ICLR, and ACL.

Qualifications

  • Research background in ML evaluation or benchmarking.
  • Expertise in LLM/frontier-model benchmarking.
  • Understanding measurement constructs and psychometrics.
  • Interest in safety and robustness of evaluation.
  • Ability to lead a team and maintain high standards.
  • Experience framing, running studies, and writing results.

Responsibilities

  • Evaluation research: convert targets into benchmark designs.
  • Benchmark development with expert-verified ground truth.
  • Recruit and calibrate a pool of experts across coding, agents, and reasoning.
  • Serve as the technical contact for labs with CEO support.
  • Turn lab requests into pilots and publish benchmarks.

Skills

ML evaluation
LLM benchmarking
Measurement science
Safety in evaluation
Team leadership
Study design

Job description

Research Scientist, LLM Evaluations & Benchmarking

Anyone AI Labs
Reports to: CEO Remote / LatAm / US

The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or measuring the wrong thing. You'll own that problem at Anyone AI, measuring frontier model capability.

This is a research role at heart: you decide what a good evaluation is , design the benchmarks that prove it, and defend the methodology under lab scrutiny. You'll build frontier-grade evaluation packages across reasoning, coding, agents, tool use, and multi-modal — grounded in expert-verified truth, validated against multiple models, and QC'd to survive buyer-side review.

Responsibilities
  • Evaluation research. Turn eval targets into original benchmark designs. Own the hard measurement questions: construct validity, item discrimination, headroom, reliability, contamination, and capability elicitation. Push toward evals that stay informative as models improve.
  • Benchmark development. Build evaluation packages with subject-matter experts, each with expert-verified ground truth, multi-model headroom results, and rigorous QC (calibration layers, severity-weighted rubrics, deterministic verifiers).
  • Experts. Recruit, calibrate, and review a pool across coding, agentic/tool-use, and STEM/reasoning. Be the final arbiter of correctness and frontier difficulty.
  • Lab relationships. Be a technical point of contact for labs, with CEO support. Understand what they're trying to measure and translate it into an evaluation design.
  • Delivery & dissemination. Turn lab requests into winning sample packages and own pilots end to end. Where the work generalizes, help turn it into public benchmarks and papers: we support publishing at venues like NeurIPS Datasets & Benchmarks, ICLR, and ACL.
What we're looking for
  • Research background in ML evaluation or benchmarking (a track record of published or open benchmarks, eval/measurement research, or equivalent hands‑on work that labs have relied on).
  • Deep LLM/frontier-model benchmarking expertise, with real strength in code-model and agentic evaluation.
  • Fluency with the measurement problem itself: construct validity, psychometrics, rubrics, pass rates, headroom, contamination, and what makes a task genuinely discriminate a model.
  • Interest in the safety side of evaluation, capability elicitation, robustness, and measuring the things that are hardest to measure honestly.
  • Proven ability to hold a team or expert pool to a rigorous standard.
  • Comfort with the full research loop: framing the question, running the study, and writing it up.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Research
Member of Technical Staff - Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Relocation support
Health insurance
Meals provided
+2
LLM Benchmarking Research Scientist
LLM Benchmarking Research Scientist

Anyone AI Inc. • Northern (KY)

Hybrid
USD 110,000 - 160,000
Head of Research
Head of Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 220,000 - 340,000
Relocation support
Health insurance
Meals provided (Lunch/Dinner)
+3
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Scale AI • San Francisco (CA)

On-site
USD 166,000 - 207,000
Health coverage
Equity
Retirement benefits
+3
Head of Research
Head of Research

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 225,000 - 275,000
Relocation and transportation support
Health and dental insurance
Lunch and dinner provided
+6
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Base salary + equity
Health, dental, vision coverage
Retirement benefits
+3
Member of Technical Staff (Language Model Evaluations)
Member of Technical Staff (Language Model Evaluations)

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Evaluations Engineer
Evaluations Engineer

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 190,000
Relocation and transportation support
Health/dental insurance coverage
Lunch and dinner provided
+2
Senior AI Model Evaluation Scientist (LLM Benchmarks)
Senior AI Model Evaluation Scientist (LLM Benchmarks)

Cohere • Seattle (WA)

On-site
USD 180,000 - 385,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K / Pension
+5