AI Harness Engineer: LLM Judge & Benchmarking

Valarian

Greater London

Hybrid

GBP 95,000 - 140,000

Full time

11 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Competitive salary
Employer pension
Private health insurance
Hybrid work
Company retreats

Job summary

Valarian Technologies Limited is seeking an experienced AI Harness Engineer to own the experimental design, evaluation methodologies, and benchmark infrastructure for AI models and autonomous workloads. You will bridge data science, statistical validation, and production agent scaffolding, collaborating across teams in a hybrid London-based setup.

The role emphasizes rigorous hypothesis testing, robust metric design, and hands-on Python tooling to advance evaluation pipelines and guardrails for

Qualifications

  • Strong grounding in statistics and experimental design.
  • Experience validating LLM judge setups against human annotations.
  • Proven ability to design evaluation datasets with precise scoring rubrics.
  • Experience evaluating non-deterministic multi-step systems.
  • Ability to define metrics tied to real-world outcomes.
  • Proficient in error analysis and communicating findings.

Responsibilities

  • Architect Evaluation Runtimes & Harnesses: Design and maintain scalable Python evaluation harnesses that measure task completion, trajectory quality, tool-use correctness, cost, latency, and variance across multi-step agentic systems.
  • Design Experiments & Validate Results: Apply rigorous statistical methods to determine whether performance changes are true improvements or stochastic noise.
  • Build & Calibrate LLM-as-a-Judge Pipelines: Develop automated judging systems validated against human ground truth; measure agreement metrics and detect judge biases.
  • Curate Benchmark Datasets & Rubrics: Define sampling strategies, detailed annotation guidelines, and scoring rubrics; manage label noise to ensure benchmark integrity.
  • Deep Error & Trajectory Analysis: Conduct hands-on failure analysis on agent runs, cluster root causes, and communicate findings to teams.
  • Drive Harness Scaffolding Iteration: Translate evaluation results into architectural improvements across agent scaffolding, prompt design, tool execution loops, context window management.

Skills

Statistics & Design
LLM Evaluation
Evaluation Dataset
Non-Deterministic Systems
Metric Design
Error Analysis
LLM Integration

Tools

Python
LangChain
AutoGen
lm-evaluation-harness
Promptfoo
Ragas

Job description

Valarian Technologies Limited is seeking an experienced AI Harness Engineer to own the experimental design, evaluation methodologies, and benchmark infrastructure for AI models and autonomous workloads. You will bridge data science, statistical validation, and production agent scaffolding, collaborating across teams in a hybrid London-based setup.

The role emphasizes rigorous hypothesis testing, robust metric design, and hands-on Python tooling to advance evaluation pipelines and guardrails for

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Harness Engineer — LLM Evaluation & Benchmark Architect
AI Harness Engineer — LLM Evaluation & Benchmark Architect

Valarian Technologies • Greater London

Hybrid
GBP 75,000 - 110,000
Equity
Competitive salary
Employer pension contributions
+3
AI Harness Engineer — Equity + Hybrid Work
AI Harness Engineer — Equity + Hybrid Work

Valarian • Greater London

Hybrid
GBP 90,000 - 130,000
Equity
Competitive salary
Employer pension
+3
AI Engineer (Harness)
AI Engineer (Harness)

Valarian Technologies • Greater London

Hybrid
GBP 75,000 - 110,000
Equity
Competitive salary
Employer pension contributions
+3
Benchmarking Engineer for Agentic AI Systems (London)
Benchmarking Engineer for Agentic AI Systems (London)

Callosum • Greater London

On-site
GBP 90,000 - 140,000
Equity & Ownership
Private healthcare
Visa sponsorship and relocation
+1
Engineering AI Engineer (Harness) London — Full time Apply →
Engineering AI Engineer (Harness) London — Full time Apply →

Valarian • Greater London

Hybrid
GBP 95,000 - 140,000
Equity
Competitive salary
Employer pension
+3
AI Engineer (Harness)
AI Engineer (Harness)

Valarian • Greater London

Hybrid
GBP 90,000 - 130,000
Equity
Competitive salary
Employer pension
+3
Director, Harness Engineering — AI Governance & Scale
Director, Harness Engineering — AI Governance & Scale

FICO • England

On-site
GBP 110,000 - 160,000
Competitive compensation
Career development
Inclusive culture
+1
Production AI Engineer: LLM Deployments & Impact
Production AI Engineer: LLM Deployments & Impact

Harrington Starr • England

On-site
GBP 90,000 - 120,000
Research Engineer — AI Agent Harness & Eval
Research Engineer — AI Agent Harness & Eval

Mat Vin • Greater London

Hybrid
GBP 120,000 - 160,000
Health insurance
Pension contributions
Hybrid work model
+2
Hybrid Agentic AI Engineer – LLM Systems
Hybrid Agentic AI Engineer – LLM Systems

Experis UK • Greater London

Hybrid
GBP 70,000 - 110,000
Hybrid work model
London office option