AI Research scientist - Evals

Cerebro

San Francisco (CA)

On-site

USD 165,000 - 195,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Relocation support
Health and dental insurance
Unlimited PTO
On-site benefits

Job summary

Cerebro invites applications for a Research Scientist to design novel benchmarks and evaluate frontier language models and agents. You will lead research, design experiments and collaborate with research engineers, foundation-model developers and domain experts to turn open-ended questions into scalable evaluations.

Unlike many industry roles, this position provides a direct route from research to external impact with opportunities to publish and present findings that influence how leading AI

Qualifications

  • Master’s degree, PhD or equivalent research experience in ML, NLP, CS or related field.
  • Strong research track record in ML, NLP, evaluation, benchmarking or an adjacent area.
  • Publications at respected conferences or journals such as NeurIPS, ICML, ICLR, ACL or EMNLP.
  • Excellent knowledge of experimental design, statistical analysis and research methodology.
  • The ability to translate open-ended research questions into precise, testable experiments.
  • Clear written and verbal communication skills.
  • A highly independent approach and the ability to operate effectively within a small, research-intensive team.

Responsibilities

  • Design novel benchmarks for evaluating frontier language models and agents.
  • Research the limitations of existing evaluation methodologies, including benchmark contamination, model-based grading and reward hacking.
  • Develop experiments that distinguish genuine capability improvements from memorisation or benchmark optimisation.
  • Evaluate models across realistic, long-horizon and economically valuable tasks.
  • Analyse model performance, failure modes and emergent behaviours.
  • Establish statistically rigorous methods for validating benchmark quality and reliability.
  • Collaborate with research engineers to implement and run evaluations at scale.
  • Work with external research and industry partners to understand emerging evaluation.

Skills

Master's/PhD in ML/CS
Experimental design
Statistical analysis
Research publications

Education

Master's/PhD in ML/CS

Job description

Research Scientist – Frontier AI Evaluations
Compensation: $165,000–$195,000 base salary + equity

Cerebro Partners is working with a fast-growing AI research company developing rigorous benchmarks and evaluation infrastructure for frontier language models and agents.

As AI systems become increasingly capable, many established benchmarks are becoming saturated, contaminated or disconnected from the tasks these models are expected to perform in the real world. This team is building the evaluation methods needed to measure genuine progress—and establish where frontier systems remain unreliable.

They are now hiring a Research Scientist to design novel benchmarks, study emerging model capabilities and help advance the science of AI evaluation.

The Role

You will lead research into how frontier models and agents should be evaluated across complex, economically valuable and long-horizon tasks.

Your work will combine original research, experimental design and practical implementation. You will collaborate with research engineers, foundation-model developers and domain experts to turn open-ended questions about model behaviour into rigorous, scalable evaluations.

Unlike many industrial research positions, this role provides a direct route from research to external impact. You will have opportunities to publish, present your findings and see your work influence how leading AI systems are developed and deployed.

What You’ll Do
  • Design novel benchmarks for evaluating frontier language models and agents
  • Research the limitations of existing evaluation methodologies, including benchmark contamination, model-based grading and reward hacking
  • Develop experiments that distinguish genuine capability improvements from memorisation or benchmark optimisation
  • Evaluate models across realistic, long-horizon and economically valuable tasks
  • Analyse model performance, failure modes and emergent behaviours
  • Establish statistically rigorous methods for validating benchmark quality and reliability
  • Collaborate with research engineers to implement and run evaluations at scale
  • Work with external research and industry partners to understand emerging evaluation
requirements
  • Publish research findings and contribute to the wider AI evaluation community
What We’re Looking For
  • A Master’s degree, PhD or equivalent research experience in machine learning, natural language processing, computer science or a related field
  • A strong record of research in machine learning, NLP, evaluation, benchmarking or an adjacent area
  • Publications at respected conferences or journals such as NeurIPS, ICML, ICLR, ACL or EMNLP
  • Excellent knowledge of experimental design, statistical analysis and research methodology
  • The ability to translate open-ended research questions into precise, testable experiments
  • Clear written and verbal communication skills
  • A highly independent approach and the ability to operate effectively within a small, research-intensive team
Particularly Relevant Experience
  • Designing LLM benchmarks, datasets or evaluation frameworks
  • Evaluating agents, tool use, reasoning or long-context capabilities
  • Working with human or model-based evaluation methods
  • Research into benchmark validity, contamination, robustness or reward hacking
  • Building research systems or running experiments at scale
  • Applied research experience within an AI lab or early-stage technology company
  • Contributions to open-source evaluation tools or datasets
Why Join?
  • Work on one of the defining research problems in frontier AI
  • Develop evaluations used to understand newly released models and emerging capabilities
  • Collaborate with leading foundation-model labs and technical organizations
  • Maintain the opportunity to publish and present your research
  • Join a small, highly technical team where individual researchers have significant ownership
  • Receive meaningful equity alongside a competitive base salary
  • Access relocation support, comprehensive health and dental insurance, unlimited PTO, and additional on-site benefits

This is an on-site position in San Francisco. Candidates relocating from elsewhere in the United States or internationally are encouraged to apply.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff (Language Model Evaluations)
Member of Technical Staff (Language Model Evaluations)

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 170,000 - 230,000
Equity
Frontier AI exposure
Frontier AI Evaluation Scientist (Equity + Relocation)
Frontier AI Evaluation Scientist (Equity + Relocation)

Cerebro • San Francisco (CA)

On-site
USD 165,000 - 195,000
Equity
Relocation support
Health and dental insurance
+2
Applied AI Researcher
Applied AI Researcher

Morpheus Talent Solutions • United States

On-site
USD 120,000 - 180,000
Member of Technical Staff - Research
Member of Technical Staff - Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Relocation support
Health insurance
Meals provided
+2
Research Scientist (Remote/US/LATAM)
Research Scientist (Remote/US/LATAM)

Anyone AI • United States

Remote
MXN 2,626,000 - 3,678,000
Member of Technical Staff (Language Model Evaluations)
Member of Technical Staff (Language Model Evaluations)

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Research Engineer — AI Alignment & Evaluation
Research Engineer — AI Alignment & Evaluation

W3 Sourcing • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Research Scientist, Frontier Risk Evaluations
Research Scientist, Frontier Risk Evaluations

United States Digital Space LLC • New York (NY), San Francisco (CA)

On-site
USD 216,000 - 270,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
Senior AI Forward Deployed Engineer
Senior AI Forward Deployed Engineer

Handshake • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Equity in a fast-growing company
401(k) match
Paid parental leave
+2
Member of Technical Staff, Frontier Evals
Member of Technical Staff, Frontier Evals

Intelligence • San Francisco (CA)

On-site
USD 180,000 - 280,000
Meaningful equity