AI Harness Engineer — LLM Evaluation & Benchmark Architect

Valarian Technologies

Greater London

Hybrid

GBP 75,000 - 110,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Competitive salary
Employer pension contributions
Private health insurance
Hybrid work setup
Company retreats and meetups

Job summary

Valarian Technologies is seeking an AI Harness Engineer to own experimental design and evaluation infrastructure for AI models and autonomous workloads. You will bridge data science, statistics, and production scaffolds, designing robust datasets, calibrated LLM-judge pipelines, and scalable Python harnesses.

You will drive metric integrity, error analysis, and governance of tool orchestration, ensuring safe, reproducible evaluations while collaborating with London-based teams in a hybrid setup.

Qualifications

  • Demonstrated expertise in hypothesis testing, bootstrapping, power analysis, and statistical significance within probabilistic ML systems.
  • Proven experience building, auditing, and validating LLM judge setups against human annotations, including agreement metrics and bias mitigation.
  • Strong track record designing evaluation datasets, crafting scoring rubrics, managing label noise, and measuring inter-annotator agreement.
  • Deep comfort evaluating complex agent trajectories, tool execution sequences, latency, cost, and variance across repeated runs, not just single-shot benchmarks.
  • Exceptional judgment in defining metrics tied to real-world outcomes with attention to benchmark gaming, drift, or shortcut learning.
  • Ability to dissect complex execution logs, categorize error modes, and communicate data-driven recommendations.

Responsibilities

  • Architect Evaluation Runtimes & Harnesses: design scalable Python evaluation harnesses measuring task completion, trajectory quality, tool-use correctness, cost, latency, and variance.
  • Design Experiments & Validate Results: apply statistical methods to determine true improvements vs. noise.
  • Build & Calibrate LLM-as-a-Judge Pipelines: develop automated judging systems with human-ground-truth validation and bias detection.
  • Curate Benchmark Datasets & Rubrics: define sampling, annotation guidelines, and scoring rubrics with inter-annotator agreement checks.
  • Deep Error & Trajectory Analysis: perform failure analysis and communicate findings to teams.
  • Drive Harness Scaffolding Iteration: translate results into architectural improvements across scaffolding, prompts, tool loops, and guardrails.

Skills

Statistics & experimental design
LLM judge validation
Evaluation dataset engineering
Non-deterministic evaluation
Metric design
Error analysis
LLM integration
Agent architecture
Harness development

Tools

LangChain
LangGraph
AutoGen
lm-evaluation-harness
Promptfoo
Ragas
Kubernetes

Job description

Valarian Technologies is seeking an AI Harness Engineer to own experimental design and evaluation infrastructure for AI models and autonomous workloads. You will bridge data science, statistics, and production scaffolds, designing robust datasets, calibrated LLM-judge pipelines, and scalable Python harnesses.

You will drive metric integrity, error analysis, and governance of tool orchestration, ensuring safe, reproducible evaluations while collaborating with London-based teams in a hybrid setup.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Harness Engineer: LLM Judge & Benchmarking
AI Harness Engineer: LLM Judge & Benchmarking

Valarian • Greater London

Hybrid
GBP 95,000 - 140,000
Equity
Competitive salary
Employer pension
+3
AI Harness Engineer — Equity + Hybrid Work
AI Harness Engineer — Equity + Hybrid Work

Valarian • Greater London

Hybrid
GBP 90,000 - 130,000
Equity
Competitive salary
Employer pension
+3
AI Engineer (Harness)
AI Engineer (Harness)

Valarian Technologies • Greater London

Hybrid
GBP 75,000 - 110,000
Equity
Competitive salary
Employer pension contributions
+3
Engineering AI Engineer (Harness) London — Full time Apply →
Engineering AI Engineer (Harness) London — Full time Apply →

Valarian • Greater London

Hybrid
GBP 95,000 - 140,000
Equity
Competitive salary
Employer pension
+3
AI Engineer (Harness)
AI Engineer (Harness)

Valarian • Greater London

Hybrid
GBP 90,000 - 130,000
Equity
Competitive salary
Employer pension
+3
Director, Harness Engineering — AI Governance & Scale
Director, Harness Engineering — AI Governance & Scale

FICO • England

On-site
GBP 110,000 - 160,000
Competitive compensation
Career development
Inclusive culture
+1
Production AI Engineer: LLM Deployments & Impact
Production AI Engineer: LLM Deployments & Impact

Harrington Starr • England

On-site
GBP 90,000 - 120,000
Benchmarking Engineer for Agentic AI Systems (London)
Benchmarking Engineer for Agentic AI Systems (London)

Callosum • Greater London

On-site
GBP 90,000 - 140,000
Equity & Ownership
Private healthcare
Visa sponsorship and relocation
+1
Hybrid Agentic AI Engineer – LLM Systems
Hybrid Agentic AI Engineer – LLM Systems

Experis UK • Greater London

Hybrid
GBP 70,000 - 110,000
Hybrid work model
London office option
Research Engineer — AI Agent Harness & Eval
Research Engineer — AI Agent Harness & Eval

Mat Vin • Greater London

Hybrid
GBP 120,000 - 160,000
Health insurance
Pension contributions
Hybrid work model
+2