AI Benchmarking Architect

AI Startups UK

Greater London

Hybrid

GBP 120,000 - 180,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive salary
Equity & ownership
Private healthcare
Visa sponsorship and relocation
London office

Job summary

Callosum is building a unified benchmarking system for evaluating agentic and algorithmic AI work. The role focuses on reproducible, contamination‑controlled evaluation across motifs and decomposition strategies, grounded in real execution traces and sandboxed grading.

You will architect and implement pipelines that measure task success, quality and robustness, influencing which approaches ship and how we communicate proof points to customers.

Qualifications

  • PhD or equivalent in computer science, ML, or related field.
  • Authorship or co-authorship of benchmarks or evaluation papers at NeurIPS/ICML/ICLR/ACL.
  • Strong understanding of evaluation pitfalls: contamination, overfitting, irreproducibility.
  • Engineering ability: Python expertise, sandboxed/distributed execution, CI.
  • Hands-on experience evaluating agentic or multi-step LLM systems.

Responsibilities

  • Design and build a unified system for evaluating agentic and algorithmic solutions across workloads, tracking task success, quality and robustness.
  • Mine real commits and traces, run sandboxed execution grading, and build task suites reflecting real agentic work.
  • Enforce controls against contamination, overfitting, and metric gaming; keep baselines stable and reproducible.
  • Compare algorithmic and agentic approaches; determine which to ship based on resolved-task quality.
  • Lead external benchmark co-publications and ensure alignment with peer and customer review.
  • Incorporate results into product decisions and customer communications.

Job description

Callosum is building a unified benchmarking system for evaluating agentic and algorithmic AI work. The role focuses on reproducible, contamination‑controlled evaluation across motifs and decomposition strategies, grounded in real execution traces and sandboxed grading.

You will architect and implement pipelines that measure task success, quality and robustness, influencing which approaches ship and how we communicate proof points to customers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Benchmarking Engineer for Agentic AI Systems (London)
Benchmarking Engineer for Agentic AI Systems (London)

Callosum • Greater London

On-site
GBP 90,000 - 140,000
Equity & Ownership
Private healthcare
Visa sponsorship and relocation
+1
AI Benchmarking Engineer for Cross-Stack Evaluation
AI Benchmarking Engineer for Cross-Stack Evaluation

Callosum Technologies Ltd. • Greater London

Hybrid
GBP 75,000 - 120,000
Visa sponsorship
Relocation benefits
Staff Research Engineer: Agentic Evaluation & Benchmarks
Staff Research Engineer: Agentic Evaluation & Benchmarks

Callosum Technologies Ltd. • Greater London

Hybrid
GBP 90,000 - 140,000
Competitive salary
Equity & Ownership
Private healthcare
+2
Agentic Evaluation Engineer — Benchmarks & Systems
Agentic Evaluation Engineer — Benchmarks & Systems

AI Startups UK • Greater London

Hybrid
GBP 90,000 - 130,000
Competitive salary
Equity ownership
Private healthcare
+1
Agentic AI Research Engineer — Systems & Evaluation
Agentic AI Research Engineer — Systems & Evaluation

Callosum • Greater London

On-site
GBP 120,000 - 190,000
Visa sponsorship
Relocation benefits
Private healthcare
+1
Agentic Evaluation Engineer, Research Scientist
Agentic Evaluation Engineer, Research Scientist

Callosum • Greater London

On-site
GBP 120,000 - 180,000
Private healthcare
Visa sponsorship
Relocation benefits
+1
AI Benchmark Scientist: Biochemistry & Genetics
AI Benchmark Scientist: Biochemistry & Genetics

Obsidian • Greater London

Remote
GBP 40,000 - 50,000
AI Benchmark Scientist: Biology PhD Coder
AI Benchmark Scientist: Biology PhD Coder

Mercor • Greater London

On-site
GBP 55,000 - 83,000
Senior AI Systems Engineer — Scale & Benchmarking
Senior AI Systems Engineer — Scale & Benchmarking

Graphcore • West of England

On-site
GBP 70,000 - 110,000
Flexible benefits
Generous leave
Pension matching
+2
Benchmark Designer in Computational Statistics & Applied Math
Benchmark Designer in Computational Statistics & Applied Math

Obsidian • Greater London

On-site
GBP 45,000 - 75,000