Frontier AI Evaluation Scientist (Equity + Relocation)

Cerebro

San Francisco (CA)

On-site

USD 165,000 - 195,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Relocation support
Health and dental insurance
Unlimited PTO
On-site benefits

Job summary

Cerebro invites applications for a Research Scientist to design novel benchmarks and evaluate frontier language models and agents. You will lead research, design experiments and collaborate with research engineers, foundation-model developers and domain experts to turn open-ended questions into scalable evaluations.

Unlike many industry roles, this position provides a direct route from research to external impact with opportunities to publish and present findings that influence how leading AI

Qualifications

  • Master’s degree, PhD or equivalent research experience in ML, NLP, CS or related field.
  • Strong research track record in ML, NLP, evaluation, benchmarking or an adjacent area.
  • Publications at respected conferences or journals such as NeurIPS, ICML, ICLR, ACL or EMNLP.
  • Excellent knowledge of experimental design, statistical analysis and research methodology.
  • The ability to translate open-ended research questions into precise, testable experiments.
  • Clear written and verbal communication skills.
  • A highly independent approach and the ability to operate effectively within a small, research-intensive team.

Responsibilities

  • Design novel benchmarks for evaluating frontier language models and agents.
  • Research the limitations of existing evaluation methodologies, including benchmark contamination, model-based grading and reward hacking.
  • Develop experiments that distinguish genuine capability improvements from memorisation or benchmark optimisation.
  • Evaluate models across realistic, long-horizon and economically valuable tasks.
  • Analyse model performance, failure modes and emergent behaviours.
  • Establish statistically rigorous methods for validating benchmark quality and reliability.
  • Collaborate with research engineers to implement and run evaluations at scale.
  • Work with external research and industry partners to understand emerging evaluation.

Skills

Master's/PhD in ML/CS
Experimental design
Statistical analysis
Research publications

Education

Master's/PhD in ML/CS

Job description

Cerebro invites applications for a Research Scientist to design novel benchmarks and evaluate frontier language models and agents. You will lead research, design experiments and collaborate with research engineers, foundation-model developers and domain experts to turn open-ended questions into scalable evaluations.

Unlike many industry roles, this position provides a direct route from research to external impact with opportunities to publish and present findings that influence how leading AI

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Research scientist - Evals
AI Research scientist - Evals

Cerebro • San Francisco (CA)

On-site
USD 165,000 - 195,000
Equity
Relocation support
Health and dental insurance
+2
Frontier Language Model Evaluation Engineer
Frontier Language Model Evaluation Engineer

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 170,000 - 230,000
Equity
Frontier AI exposure
Staff AI Research Engineer — Frontier Model Evaluations
Staff AI Research Engineer — Frontier Model Evaluations

Intelligence • San Francisco (CA)

On-site
USD 180,000 - 280,000
Meaningful equity
Research Engineer - Frontier AI Training & Evals
Research Engineer - Frontier AI Training & Evals

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Competitive compensation
Medical, dental, vision coverage
Lunch and dinner in office
+5
Frontier AI Evaluation Architect
Frontier AI Evaluation Architect

United States Digital Space LLC • San Francisco (CA)

On-site
USD 120,000 - 160,000
Applied Research Scientist — Frontier AI (Equity)
Applied Research Scientist — Frontier AI (Equity)

Fleet AI, Inc. • Buffalo (NY)

On-site
USD 150,000 - 210,000
Frontier AI Evaluation Engineer: RL Environments
Frontier AI Evaluation Engineer: RL Environments

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Frontier AI Risk Evaluation Scientist
Frontier AI Risk Evaluation Scientist

Scale • San Francisco (CA), New York (NY)

On-site
USD 216,000 - 270,000
Health coverage
Retirement benefits
Learning stipend
+2
Frontier RL Evaluation Engineer
Frontier RL Evaluation Engineer

Key Talent Solutions • San Francisco (CA)

On-site
USD 180,000 - 350,000
Frontier AI Risk Evaluation Scientist
Frontier AI Risk Evaluation Scientist

United States Digital Space LLC • New York (NY), San Francisco (CA)

On-site
USD 216,000 - 270,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2