AI Engineer (Harness)

Valarian Technologies

Greater London

Hybrid

GBP 75,000 - 110,000

Full time

10 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Competitive salary
Employer pension contributions
Private health insurance
Hybrid work setup
Company retreats and meetups

Job summary

Valarian Technologies is seeking an AI Harness Engineer to own experimental design and evaluation infrastructure for AI models and autonomous workloads. You will bridge data science, statistics, and production scaffolds, designing robust datasets, calibrated LLM-judge pipelines, and scalable Python harnesses.

You will drive metric integrity, error analysis, and governance of tool orchestration, ensuring safe, reproducible evaluations while collaborating with London-based teams in a hybrid setup.

Qualifications

  • Demonstrated expertise in hypothesis testing, bootstrapping, power analysis, and statistical significance within probabilistic ML systems.
  • Proven experience building, auditing, and validating LLM judge setups against human annotations, including agreement metrics and bias mitigation.
  • Strong track record designing evaluation datasets, crafting scoring rubrics, managing label noise, and measuring inter-annotator agreement.
  • Deep comfort evaluating complex agent trajectories, tool execution sequences, latency, cost, and variance across repeated runs, not just single-shot benchmarks.
  • Exceptional judgment in defining metrics tied to real-world outcomes with attention to benchmark gaming, drift, or shortcut learning.
  • Ability to dissect complex execution logs, categorize error modes, and communicate data-driven recommendations.

Responsibilities

  • Architect Evaluation Runtimes & Harnesses: design scalable Python evaluation harnesses measuring task completion, trajectory quality, tool-use correctness, cost, latency, and variance.
  • Design Experiments & Validate Results: apply statistical methods to determine true improvements vs. noise.
  • Build & Calibrate LLM-as-a-Judge Pipelines: develop automated judging systems with human-ground-truth validation and bias detection.
  • Curate Benchmark Datasets & Rubrics: define sampling, annotation guidelines, and scoring rubrics with inter-annotator agreement checks.
  • Deep Error & Trajectory Analysis: perform failure analysis and communicate findings to teams.
  • Drive Harness Scaffolding Iteration: translate results into architectural improvements across scaffolding, prompts, tool loops, and guardrails.

Skills

Statistics & experimental design
LLM judge validation
Evaluation dataset engineering
Non-deterministic evaluation
Metric design
Error analysis
LLM integration
Agent architecture
Harness development

Tools

LangChain
LangGraph
AutoGen
lm-evaluation-harness
Promptfoo
Ragas
Kubernetes

Job description

Valarian Technologies is a dual-use technology company building critical tools to safeguard the future in an era of evolving global security challenges. We're rethinking security beyond traditional military domains, addressing asymmetric threats that impact our technological advantage, economic strength, and democratic institutions.

We build Acra – the platform foundation for everything we do as a dual-use technology company. The platform’s name, rooted in the Greek word for citadel (or, fortress), reflects the design and purpose of our infrastructure-agnostic secure enclaves: protecting critical data. Some of the government and commercial workflows include: increased operational resiliency for mission-critical systems and functions; enabling organisations to more quickly and widely adopt emerging technologies while ensuring the integrity of their intellectual property; information flow during disaster response scenarios, and zero-trust / least-privilege environments for M&A, attorney-client privileged communications, etc. And we’ve only scratched the surface.

At our core, we're driven by a shared mission and a belief in making a tangible impact on our world. Whether you join our London HQ or the wider global organisation, you’ll be a part of collaborative, high-performing teams, creating cutting‑edge software, platforms, and infrastructure.

The Role

As an AI Harness Engineer at Valarian, you will own the experimental design, evaluation methodologies, and benchmark infrastructure for our AI models and autonomous agentic workloads. In an emerging domain with no standard playbook, you will serve as the bridge between rigorous data science, statistical validation, and production agent scaffolding.

You will design robust evaluation datasets, architect calibrated LLM-as-a-judge pipelines, and build metrics that accurately capture multi-step agent performance under non-deterministic conditions. Your insights and error analyses will directly drive iterative improvements to our harness scaffolding, tool orchestration, and system safety guardrails.

What you’ll do:
  • Architect Evaluation Runtimes & Harnesses: Design and maintain scalable Python evaluation harnesses that measure task completion, trajectory quality, tool-use correctness, cost, latency, and variance across multi-step agentic systems.

  • Design Experiments & Validate Results: Apply rigorous statistical methods—including hypothesis testing, confidence intervals, bootstrapping, and power analysis—to determine whether performance changes are true system improvements or stochastic noise.

  • Build & Calibrate LLM-as-a-Judge Pipelines: Develop automated judging systems validated against human ground truth. Measure alignment using agreement metrics (Cohen’s kappa, correlation, precision/recall) and systematically detect judge failure modes such as position bias, verbosity bias, self-preference, and prompt sensitivity.

  • Curate Benchmark Datasets & Rubrics: Define sampling strategies, detailed annotation guidelines, and scoring rubrics. Measure inter‑annotator agreement and manage label noise to ensure benchmark integrity over time.

  • Deep Error & Trajectory Analysis: Conduct hands‑on failure analysis on agent runs, cluster root causes into actionable error categories, and clearly communicate findings to both engineering teams and stakeholders.

  • Drive Harness Scaffolding Iteration: Translate evaluation results directly into architectural improvements across agent scaffolding, prompt design, tool execution loops, context window compaction, retries, and safety guardrails.

What we are looking for:
  • Grounding in Statistics & Experimental Design: Demonstrated expertise in hypothesis testing, bootstrapping, power analysis, and statistical significance within probabilistic machine learning systems.

  • LLM-as-a-Judge Validation: Proven experience building, auditing, and validating LLM judge setups against human annotations, including tracking agreement metrics and mitigating judge biases.

  • Evaluation Dataset Engineering: Strong track record designing evaluation datasets, crafting precise scoring rubrics, managing label noise, and measuring inter‑annotator agreement.

  • Evaluation of Non‑Deterministic Multi‑Step Systems: Deep comfort evaluating complex agent trajectories, tool execution sequences, latency, cost, and variance across repeated runs rather than relying solely on single‑shot benchmarks.

  • Metric Design & Integrity: Exceptional judgment in defining metrics tied to real‑world outcomes, with a keen eye for detecting benchmark gaming, target drift, or shortcut learning.

  • Rigor in Error Analysis: Ability to dissect complex execution logs, categorize failure modes, and synthesize clear, data‑driven recommendations.

  • Hands‑On LLM Integration: Practical proficiency with LLM APIs, prompt engineering, structured outputs, function/tool calling, and retrieval systems in Python.

  • Agent Architecture Awareness: Clear understanding of planning loops, tool orchestration, memory management, and multi‑agent coordination, including their trade‑offs and common failure modes.

  • Harness & Scaffolding Contribution: Ability to write clean, maintainable Python code to iterate on harness components such as guardrails, retry logic, state handling, and context management.

Nice to have:
  • Experience with agent and evaluation frameworks (e.g., LangChain, LangGraph, AutoGen, lm-evaluation-harness, Promptfoo, or Ragas).
  • Exposure to running evaluation pipelines or agent workloads within secure, enclave, or Kubernetes environments.
Benefits:

Our benefits are designed to ensure our employees feel taken care of and are proud to be a part of the Valarian team. We are committed to consistently enhancing our benefit package, taking into account the overall well‑being and needs of our teammates. Here are the key benefits accessible to all employees at Valarian Technologies:

  • Equity – because you have the right to own what you’re building
  • A competitive salary – because we value your unique skills
  • Employer pension contributions – because you deserve a secure future
  • Private health insurance - because your health is important to us
  • Hybrid work setup – because everyone has different needs
  • Rewarding company retreats and meetups that respect your work/life balance – because we love getting to know each other!
Life at Valarian

Our culture is built on inclusivity, compassion and flexibility – we want everyone to be empowered to achieve their goals at Valarian.

The work we do is vital, but so are the connections that make it happen. We thrive on the shared energy, spontaneous conversations, and mutual trust built when we spend time together. We operate on a hybrid model, gathering in our London office 3 days a week to support one another and collaborate.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Valarian Technologies Limited is an equal opportunity employer and welcomes applications from individuals regardless of race, colour, religion, sex, sexual orientation, gender, identity or expression, national origin, age, disability, genetic information, marital status, veteran, amnesty, or any other legally protected characteristic.

We are committed to ensuring a fair and inclusive recruitment process and providing employment opportunities to all applicants. Decision recruitment, hiring, and employment are based solely on qualifications, skills, and experience relevant to the job requirements.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Engineering AI Engineer (Harness) London — Full time Apply →
Engineering AI Engineer (Harness) London — Full time Apply →

Valarian • Greater London

Hybrid
GBP 95,000 - 140,000
Equity
Competitive salary
Employer pension
+3
AI Engineer (Harness)
AI Engineer (Harness)

Valarian • Greater London

Hybrid
GBP 90,000 - 130,000
Equity
Competitive salary
Employer pension
+3
Software Engineer
Software Engineer

Valarian Technologies • Greater London

Hybrid
GBP 80,000 - 120,000
Equity
Competitive salary
Employer pension contributions
+3
Senior Technical Product Manager
Senior Technical Product Manager

Valarian • Greater London

On-site
GBP 90,000 - 120,000
Equity
Competitive salary
Employer pension contributions
+3
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Valarian Technologies • Greater London

Hybrid
GBP 90,000 - 130,000
Equity
Competitive salary
Employer pension contributions
+3
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Valarian • Greater London

Hybrid
GBP 95,000 - 150,000
Equity
Salary
Employer pension contributions
+3
Senior Software Engineer
Senior Software Engineer

Valarian Technologies • Greater London

Hybrid
GBP 90,000 - 160,000
Equity
Competitive salary
Employer pension contributions
+3
Platform Engineer
Platform Engineer

Valarian • Greater London

Hybrid
GBP 90,000 - 125,000
Equity
Competitive salary
Employer pension contributions
+3
Release Engineer
Release Engineer

Valarian Technologies Limited • Greater London

Hybrid
GBP 95,000 - 135,000
Equity
Private health insurance
Hybrid work setup
+2
TechOps Engineer
TechOps Engineer

Valarian Technologies • Greater London

Hybrid
GBP 50,000 - 70,000
Equity
Competitive salary
Employer pension contributions
+2