AI Research Scientist

OxSci

Greater London

On-site

GBP 90,000 - 130,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Founding equity
Direct line to founders
Open research questions
Generous token budget

Job summary

OxSci in London seeks a PhD-level researcher to lead AI-reviewer evaluation, shaping how AI reviewers are judged in science.

You will design meta-evaluations, run large-scale expert-annotation studies, and build a living taxonomy of AI-reviewer failure modes. Join a founding team with equity, flexible hours, and open research questions as you push the frontier of AI-enabled peer review.

Qualifications

  • PhD or near completion in CS/ML/NLP or related field.
  • Strong Python and hands-on experience with LLM evaluation and tooling.
  • Experience designing expert-annotation studies and benchmarks.
  • Fluency in evaluation methodology and statistics.

Responsibilities

  • Own the research agenda on AI-reviewer evaluation.
  • Design meta-evaluations exposing weaknesses beyond verdicts.
  • Run expert-annotation studies at scale with robust protocols.
  • Build a living taxonomy of AI-reviewer failure modes.
  • Calibrate quality scoring by fusing human and AI judgments.
  • Translate findings into concrete improvements to review agents.

Skills

Python
LLM evaluation
Evaluation methodology
Inter-annotator agreement
RAG pipelines

Education

PhD in CS/ML/NLP or related field

Job description

Overview

OxSci is building a credit rating agency for science: a certification layer that combines AI with expert peer review, so researchers, institutions, and AI developers can assess research quality quickly and at scale. Webre early, focused, and well-resourced.

The open scientific question at the heart of this company is yours to own: where does the frontier of AI peer review actually stand? What can AI reviewers catch that human experts miss, what do they still get wrong, and how do you prove it rigorously? The standards we set now, including what \"quality\" even means and how we demonstrate our AI reviewers are actually good, will define both the company and, we believe, the field.

The tech side is led by OxSci's cofounder, a senior tech lead from one of the world\'s largest technology companies, with deep experience building and operating systems at global scale. You\'d work with both founders day to day, shaping the evaluation methodology and research culture from the ground up.

What you\'ll do
  • Own the research agenda on AI-reviewer evaluation. Track the frontier (AI-scientist, automated-review, LLM-as-a-judge, and scholarly-NLP literature), position our system against it, and decide what we measure next and why.
  • Design meta-evaluations that expose weaknesses, not just measure agreement. Build fine-grained, criticism-level evaluations of AI review agents (correctness, factual grounding, significance, sufficiency of evidence, hallucination rate, and venue/journal matching) that reveal where and why they fail, going beyond verdict-matching.
  • Run expert-annotation studies at scale. Design the protocols, rubrics, inter-annotator agreement, and statistics needed to compare AI and human reviewers credibly, including head-to-head evaluations against other AI review systems, and defend the numbers to a skeptical scientific audience.
  • Build a living taxonomy of AI-reviewer failure modes such as subfield blind spots, long-context degradation, over-anchoring, and spurious criticism, and turn each into a regression benchmark that guards against backsliding as models and prompts change.
  • Calibrate the combined rating. Define quality-scoring rubrics for human review reports and calibrate how expert and AI judgment fuse into a single, defensible rating: the core of what universities and publishers buy from us.
  • Close the loop. Translate benchmark findings into concrete improvements to our review agents (retrieval, context engineering, orchestration, model choice) and prove the gains with the same rigor you used to find the gaps.
What we\'re looking for
  • A PhD (or near completion) in CS, ML, NLP, or a related field, or an equivalent research track record. You\'re likely already working on LLM evaluation, LLM-as-a-judge, AI for science, automated peer review, or a nearby frontier.
  • A fascination with the boundary between AI and human reviewers. What each catches that the other misses, and conviction that mapping it rigorously is how trustworthy peer review gets built.
  • A track record of rigorous evaluation of ML/LLM systems: benchmark or eval-framework design, evaluator/judge models, expert-annotation study design, hallucination and factuality measurement, uncertainty quantification, or RAG evaluation. Bonus if you\'ve built evaluations that score individual criticisms rather than just verdicts.
  • Fluency in evaluation methodology and statistics: sampling, inter-annotator agreement, significance testing, and the discipline to distinguish a real effect from a lucky prompt.
  • Strong Python and hands-on habits. You build the eval harnesses and pipelines yourself, not just spec them, with enough LLM-application fluency (RAG, tool calling, orchestration) to turn a finding into a shipped improvement.
  • Bonus: publications in NLP/ML evaluation or automated peer review; open-source benchmarks or evaluator models the community actually uses; experience with scholarly content at scale.
What we offer
  • Founding seat with meaningful equity and a direct line to the founders
  • Ownership of a genuinely open research question, with encouragement to publish and present the work
  • A standing expert-reviewer network as your annotation infrastructure
  • A proprietary, growing dataset of paired human and AI reviews of real submissions
  • The rare chance to define the standard by which AI reviewers themselves are judged
  • Genuinely competitive pay; equity discussed openly
  • Generous LLM token budget for your daily work
  • Flexible working hours; fast personal growth with broad ownership from day one
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Research Scientist - Frontier AI Peer Review
AI Research Scientist - Frontier AI Peer Review

OxSci • Greater London

On-site
GBP 90,000 - 130,000
Founding equity
Direct line to founders
Open research questions
+1
Research Engineer, Benchmarking - Member of Technical Staff
Research Engineer, Benchmarking - Member of Technical Staff

Callosum • Greater London

On-site
GBP 120,000 - 180,000
Visa sponsorship
Relocation benefits
London office onsite
+3
Senior Applied AI Engineer
Senior Applied AI Engineer

SoCode Recruitment • Greater London

Hybrid
GBP 90,000 - 150,000
Equity package
Hybrid working – 2 days per week in <b
AI Engineer
AI Engineer

Solve Intelligence • Greater London

On-site
GBP 65,000 - 85,000
Competitive Salary + Significant Equity
Full visa sponsorship
Private medical insurance
+2
Research Scientist/Engineer (Science of Scheming)
Research Scientist/Engineer (Science of Scheming)

COL Limited • Greater London

On-site
GBP 100,000 - 200,000
Market competitive salary
Equity options
Flexible work hours
+4
Research Engineer, Benchmarking - Member of Technical Staff
Research Engineer, Benchmarking - Member of Technical Staff

Callosum Technologies Ltd. • Greater London

Hybrid
GBP 75,000 - 120,000
Visa sponsorship
Relocation benefits
Research Engineer, Benchmarking - Member of Technical Staff
Research Engineer, Benchmarking - Member of Technical Staff

AI Startups UK • Greater London

Hybrid
GBP 120,000 - 180,000
Competitive salary
Equity & ownership
Private healthcare
+2
Founding AI Engineer
Founding AI Engineer

nettle • Greater London

Hybrid
GBP 50,000 - 90,000
Founding team member equity stake
Competitive salary
Comprehensive health benefits
+4
Founding AI Research Scientist
Founding AI Research Scientist

Blue Wolf Digital • Greater London

Hybrid
GBP 120,000 - 170,000
Research Engineer, Evals - Member of Technical Staff
Research Engineer, Evals - Member of Technical Staff

Callosum • Greater London

On-site
GBP 120,000 - 180,000
Private healthcare
Visa sponsorship
Relocation benefits
+1