AI Benchmarking Engineer for Cross-Stack Evaluation

Callosum Technologies Ltd.

Greater London

Hybrid

GBP 75,000 - 120,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Visa sponsorship
Relocation benefits

Job summary

Callosum Technologies Ltd. is hiring for a research-driven role in London to build and run a unified benchmarking system for evaluating agentic and algorithmic LLM workflows.

You will design tasks, collect real commits/traces, and ensure robust, contamination-free evaluation with reproducible results. The role requires a PhD or equivalent track record, strong Python, and experience with sandboxed/distributed execution.

Qualifications

  • PhD in computer science, ML, or related field, or equivalent research track record.
  • Authorship of benchmark/evaluation papers at NeurIPS/ICML/ACL is preferred.
  • Understanding of evaluation pitfalls: contamination, overfitting, irreproducible results.
  • Strong Python programming and experience with sandboxed/distributed execution and CI.
  • Hands-on experience building or evaluating multi-step LLM systems.

Responsibilities

  • Design and build a unified benchmarking system for evaluating agentic and algorithmic workloads.
  • Mine real commits/traces, run sandboxed grading, and build task suites reflecting real agentic work.
  • Enforce controls against contamination, overfitting to benchmarks, and metric gaming.
  • Compare algorithmic and agentic approaches honestly, focusing on resolved-task quality.
  • Lead external benchmark co-publications and publish results.
  • Incorporate results into which approach ships and QA quality claims.

Skills

Benchmark evaluation
Sandboxed execution
Python
Agentic/LLM eval
Publications in benchmarks

Education

PhD in computer science or related field

Tools

Python

Job description

Callosum Technologies Ltd. is hiring for a research-driven role in London to build and run a unified benchmarking system for evaluating agentic and algorithmic LLM workflows.

You will design tasks, collect real commits/traces, and ensure robust, contamination-free evaluation with reproducible results. The role requires a PhD or equivalent track record, strong Python, and experience with sandboxed/distributed execution.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Benchmarking Engineer for Agentic AI Systems (London)
Benchmarking Engineer for Agentic AI Systems (London)

Callosum • Greater London

On-site
GBP 90,000 - 140,000
Equity & Ownership
Private healthcare
Visa sponsorship and relocation
+1
AI Benchmarking Architect - Research Engineer
AI Benchmarking Architect - Research Engineer

United States Digital Space LLC • Greater London

On-site
GBP 101,000 - 192,000
Competitive salary
Equity
Private healthcare
+2
Benchmarking Systems Engineer for LLM Evaluation
Benchmarking Systems Engineer for LLM Evaluation

Callosum • Greater London

On-site
GBP 120,000 - 180,000
Visa sponsorship
Relocation benefits
London office onsite
+3
Agentic Evaluation Engineer, Research Scientist
Agentic Evaluation Engineer, Research Scientist

Callosum • Greater London

On-site
GBP 120,000 - 180,000
Private healthcare
Visa sponsorship
Relocation benefits
+1
Research Engineer, Benchmarking - Member of Technical Staff
Research Engineer, Benchmarking - Member of Technical Staff

Callosum Technologies Ltd. • Greater London

Hybrid
GBP 75,000 - 120,000
Visa sponsorship
Relocation benefits
Staff Research Engineer: Agentic Evaluation & Benchmarks
Staff Research Engineer: Agentic Evaluation & Benchmarks

Callosum Technologies Ltd. • Greater London

Hybrid
GBP 90,000 - 140,000
Competitive salary
Equity & Ownership
Private healthcare
+2
Research Engineer, Benchmarking - Member of Technical Staff
Research Engineer, Benchmarking - Member of Technical Staff

Callosum • Greater London

On-site
GBP 120,000 - 180,000
Visa sponsorship
Relocation benefits
London office onsite
+3
Research Engineer, Benchmarking - Member of Technical Staff
Research Engineer, Benchmarking - Member of Technical Staff

United States Digital Space LLC • Greater London

On-site
GBP 101,000 - 192,000
Competitive salary
Equity
Private healthcare
+2
AI Benchmark Engineer: Backend Task Design & Evaluation
AI Benchmark Engineer: Backend Task Design & Evaluation

Mindrift • Glasgow

On-site
GBP 21,000 - 35,000
Staff Engineer: Evolutionary Optimisation
Staff Engineer: Evolutionary Optimisation

Callosum Technologies Ltd. • Greater London

Hybrid
GBP 90,000 - 120,000
Equity & ownership
Private healthcare
Visa sponsorship & relocation benefits
+1