Research Scientist – Frontier AI Evaluations
Compensation: $165,000–$195,000 base salary + equity
Cerebro Partners is working with a fast-growing AI research company developing rigorous benchmarks and evaluation infrastructure for frontier language models and agents.
As AI systems become increasingly capable, many established benchmarks are becoming saturated, contaminated or disconnected from the tasks these models are expected to perform in the real world. This team is building the evaluation methods needed to measure genuine progress—and establish where frontier systems remain unreliable.
They are now hiring a Research Scientist to design novel benchmarks, study emerging model capabilities and help advance the science of AI evaluation.
The Role
You will lead research into how frontier models and agents should be evaluated across complex, economically valuable and long-horizon tasks.
Your work will combine original research, experimental design and practical implementation. You will collaborate with research engineers, foundation-model developers and domain experts to turn open-ended questions about model behaviour into rigorous, scalable evaluations.
Unlike many industrial research positions, this role provides a direct route from research to external impact. You will have opportunities to publish, present your findings and see your work influence how leading AI systems are developed and deployed.
What You’ll Do
- Design novel benchmarks for evaluating frontier language models and agents
- Research the limitations of existing evaluation methodologies, including benchmark contamination, model-based grading and reward hacking
- Develop experiments that distinguish genuine capability improvements from memorisation or benchmark optimisation
- Evaluate models across realistic, long-horizon and economically valuable tasks
- Analyse model performance, failure modes and emergent behaviours
- Establish statistically rigorous methods for validating benchmark quality and reliability
- Collaborate with research engineers to implement and run evaluations at scale
- Work with external research and industry partners to understand emerging evaluation
requirements
- Publish research findings and contribute to the wider AI evaluation community
What We’re Looking For
- A Master’s degree, PhD or equivalent research experience in machine learning, natural language processing, computer science or a related field
- A strong record of research in machine learning, NLP, evaluation, benchmarking or an adjacent area
- Publications at respected conferences or journals such as NeurIPS, ICML, ICLR, ACL or EMNLP
- Excellent knowledge of experimental design, statistical analysis and research methodology
- The ability to translate open-ended research questions into precise, testable experiments
- Clear written and verbal communication skills
- A highly independent approach and the ability to operate effectively within a small, research-intensive team
Particularly Relevant Experience
- Designing LLM benchmarks, datasets or evaluation frameworks
- Evaluating agents, tool use, reasoning or long-context capabilities
- Working with human or model-based evaluation methods
- Research into benchmark validity, contamination, robustness or reward hacking
- Building research systems or running experiments at scale
- Applied research experience within an AI lab or early-stage technology company
- Contributions to open-source evaluation tools or datasets
Why Join?
- Work on one of the defining research problems in frontier AI
- Develop evaluations used to understand newly released models and emerging capabilities
- Collaborate with leading foundation-model labs and technical organizations
- Maintain the opportunity to publish and present your research
- Join a small, highly technical team where individual researchers have significant ownership
- Receive meaningful equity alongside a competitive base salary
- Access relocation support, comprehensive health and dental insurance, unlimited PTO, and additional on-site benefits
This is an on-site position in San Francisco. Candidates relocating from elsewhere in the United States or internationally are encouraged to apply.