Benchmarking Systems Engineer for LLM Evaluation

Callosum

Greater London

On-site

GBP 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Visa sponsorship
Relocation benefits
London office onsite
Equity & ownership
Private healthcare
Salary competitive

Job summary

Callosum is seeking a PhD-level researcher to own and build a unified benchmarking system for multi-step LLM work. You will design task suites, run sandboxed grading with real commits and traces, and base decisions on reproducible results rather than model scores.

Lead external benchmark co-publications, enforce contamination controls, and drive open, scalable evaluation. This London-based, in-person role offers equity, private healthcare, visa sponsorship, and relocation support.

Qualifications

  • PhD in computer science, machine learning, or a related field, or an equivalent research track record.
  • Authorship or co-authorship of a benchmark or evaluation paper at a recognised venue – NeurIPS Datasets and Benchmarks, ICML, ICLR, ACL – ideally on agentic or LLM evaluation, or a comparably rigorous evaluation contribution.
  • A working understanding of how LLM and agent evaluation goes wrong: contamination, overfitting to benchmarks, weak baselines, underpowered comparisons, irreproducible results.
  • The engineering ability to design, build, and curate these systems decisively – strong Python, and comfort with sandboxed and distributed execution and CI.
  • Hands-on experience building or rigorously evaluating agentic or multi-step LLM systems.

Responsibilities

  • This role owns that system. You will build a harness that measures task success, quality, and robustness across motifs, agent topologies, and decomposition strategies, grounded in execution – real commits, real traces, sandboxed grading – rather than self‑reported or model‑graded scores.
  • The results become the proof points we show customers, the evidence behind the benchmarks we co‑publish, and the basis on which an approach ships or doesn’t.
  • Mine real commits and traces, run sandboxed execution grading, and build task suites that reflect real agentic work: code search, code edit and repair, repository summarisation, tool use.
  • Enforce controls against contamination, overfitting to benchmarks, and metric gaming, and keep baselines stable over time – any result should be re‑runnable to the same number, by us or by a reviewer.
  • Lead external benchmark co‑publications, held to a standard that survives peer and customer review.

Skills

Benchmarking
Agentic evaluation
Open-source tooling
Code-agent workloads
Publications

Education

PhD in CS/ML
Equivalent research track

Tools

Python
Sandboxed execution
CI/CD
Distributed execution

Job description

Callosum is seeking a PhD-level researcher to own and build a unified benchmarking system for multi-step LLM work. You will design task suites, run sandboxed grading with real commits and traces, and base decisions on reproducible results rather than model scores.

Lead external benchmark co-publications, enforce contamination controls, and drive open, scalable evaluation. This London-based, in-person role offers equity, private healthcare, visa sponsorship, and relocation support.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Benchmarking Engineer for Cross-Stack Evaluation
AI Benchmarking Engineer for Cross-Stack Evaluation

Callosum Technologies Ltd. • Greater London

Hybrid
GBP 75,000 - 120,000
Visa sponsorship
Relocation benefits
Benchmarking Engineer for Agentic AI Systems (London)
Benchmarking Engineer for Agentic AI Systems (London)

Callosum • Greater London

On-site
GBP 90,000 - 140,000
Equity & Ownership
Private healthcare
Visa sponsorship and relocation
+1
Research Engineer, Benchmarking - Member of Technical Staff
Research Engineer, Benchmarking - Member of Technical Staff

Callosum Technologies Ltd. • Greater London

Hybrid
GBP 75,000 - 120,000
Visa sponsorship
Relocation benefits
Research Engineer, Benchmarking - Member of Technical Staff
Research Engineer, Benchmarking - Member of Technical Staff

AI Startups UK • Greater London

Hybrid
GBP 120,000 - 180,000
Competitive salary
Equity & ownership
Private healthcare
+2
Research Engineer, Benchmarking - Member of Technical Staff
Research Engineer, Benchmarking - Member of Technical Staff

Callosum • Greater London

On-site
GBP 120,000 - 180,000
Visa sponsorship
Relocation benefits
London office onsite
+3
LLM Systems Engineer: Scalable AI for Science (London)
LLM Systems Engineer: Scalable AI for Science (London)

Isomorphic Labs • Greater London

Hybrid
GBP 90,000 - 130,000
Agentic Evaluation Engineer — Benchmarks & Systems
Agentic Evaluation Engineer — Benchmarks & Systems

AI Startups UK • Greater London

Hybrid
GBP 90,000 - 130,000
Competitive salary
Equity ownership
Private healthcare
+1
Staff Engineer: Evolutionary Optimisation
Staff Engineer: Evolutionary Optimisation

Callosum Technologies Ltd. • Greater London

Hybrid
GBP 90,000 - 120,000
Equity & ownership
Private healthcare
Visa sponsorship & relocation benefits
+1
Staff Engineer – LLM‑Driven Evolutionary Optimisation
Staff Engineer – LLM‑Driven Evolutionary Optimisation

AI Startups UK • Greater London

Hybrid
GBP 110,000 - 150,000
Equity & Ownership
Private healthcare
Visa sponsorship
+2
LLM-Guided Evolutionary Optimisation Engineer
LLM-Guided Evolutionary Optimisation Engineer

Callosum • Greater London

On-site
GBP 90,000 - 130,000
Equity & Ownership
Private healthcare
Visa sponsorship
+1