AI Engineer (Harness)

Valarian

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Competitive salary
Employer pension
Private health insurance
Hybrid work setup
Team retreats

Job summary

Valarian Technologies is a dual-use technology company building critical tools to safeguard the future in an era of evolving global security challenges. We’re seeking an AI Harness Engineer to own the experimental design, evaluation methodologies, and benchmark infrastructure for our AI models and autonomous agent workloads.

You will bridge data science, statistical validation, and production scaffolding to drive robust harness architecture, curate evaluation datasets, and calibrate

Qualifications

  • Expertise in hypothesis testing, bootstrapping, and power analysis within probabilistic ML systems.
  • Experience validating LLM judge setups against human annotations and tracking agreement metrics.
  • Experience designing evaluation datasets, scoring rubrics and measuring inter-annotator agreement.
  • Ability to evaluate non-deterministic multi-step agent trajectories and tool execution sequences.
  • Strong judgment in defining metrics tied to real-world outcomes and detecting benchmark gaming.

Responsibilities

  • Architect Evaluation Runtimes & Harnesses: design scalable Python evaluation harnesses for multi-step agent systems focusing on accuracy, latency and cost.
  • Design Experiments & Validate Results: apply statistical methods to determine true improvements vs stochastic noise.
  • Build & Calibrate LLM‑as‑a‑Judge Pipelines: automate judging validated against human ground truth and measure agreement metrics.
  • Curate Benchmark Datasets & Rubrics: define sampling, annotation guidelines and scoring rubrics with high benchmark integrity.
  • Deep Error & Trajectory Analysis: perform failure analysis and translate findings into actionable improvements.
  • Drive Harness Scaffolding Iteration: implement architectural improvements across agent scaffolding, prompt design, and safety guardrails.

Skills

Statistics & design
LLM evaluation
Non-deterministic eval
Python
LLM integration

Tools

LangChain

Job description

Valarian Technologies is a dual-use technology company building critical tools to safeguard the future in an era of evolving global security challenges.


We build Acra – the platform foundation for everything we do as a dual-use technology company. The platform’s name, rooted in the Greek word for citadel (or, fortress), reflects the design and purpose of our infrastructure-agnostic secure enclaves: protecting critical data. Some of the government and commercial workflows include: increased operational resiliency for mission‑critical systems and functions; enabling organisations to more quickly and widely adopt emerging technologies while ensuring the integrity of their intellectual property; information flow during disaster response scenarios, and zero‑trust / least‑privilege environments for M&A, attorney‑client privileged communications, etc. And we’ve only scratched the surface.


At our core, we're driven by a shared mission and a belief in making a tangible impact on our world. Whether you join our London HQ or the wider global organisation, you'll be a part of collaborative, high‑performing teams, creating cutting‑edge software, platforms, and infrastructure.


Valarian Technologies is a dual-use technology company building critical tools to safeguard the future in an era of evolving global security challenges. We're rethinking security beyond traditional military domains, addressing asymmetric threats that impact our technological advantage, economic strength, and democratic institutions.


We build Acra – the platform foundation for everything we do as a dual-use technology company. The platform’s name, rooted in the Greek word for citadel (or, fortress), reflects the design and purpose of our infrastructure‑agnostic secure enclaves: protecting critical data. Some of the government and commercial workflows include: increased operational resiliency for mission‑critical systems and functions; enabling organisations to more quickly and widely adopt emerging technologies while ensuring the integrity of their intellectual property; information flow during disaster response scenarios, and zero‑trust / least‑privilege environments for M&A, attorney‑client privileged communications, etc. And we’ve only scratched the surface.


At our core, we're driven by a shared mission and a belief in making a tangible impact on our world. Whether you join our London HQ or the wider global organisation, you'll be a part of collaborative, high‑performing teams, creating cutting‑edge software, platforms, and infrastructure.


The Role

As an AI Harness Engineer at Valarian, you will own the experimental design, evaluation methodologies, and benchmark infrastructure for our AI models and autonomous agentic workloads. In an emerging domain with no standard playbook, you will serve as the bridge between rigorous data science, statistical validation, and production agent scaffolding.


You will design robust evaluation datasets, architect calibrated LLM‑as‑a‑judge pipelines, and build metrics that accurately capture multi‑step agent performance under non‑deterministic conditions. Your insights and error analyses will directly drive iterative improvements to our harness scaffolding, tool orchestration, and system safety guardrails.


What you’ll do:


  • Architect Evaluation Runtimes & Harnesses: Design and maintain scalable Python evaluation harnesses that measure task completion, trajectory quality, tool‑use correctness, cost, latency, and variance across multi‑step agentic systems.

  • Design Experiments & Validate Results: Apply rigorous statistical methods—including hypothesis testing, confidence intervals, bootstrapping, and power analysis—to determine whether performance changes are true system improvements or stochastic noise.

  • Build & Calibrate LLM‑as‑a‑Judge Pipelines: Develop automated judging systems validated against human ground truth. Measure alignment using agreement metrics (Cohen's kappa, correlation, precision/recall) and systematically detect judge failure modes such as position bias, verbosity bias, self‑preference, and prompt sensitivity.

  • Curate Benchmark Datasets & Rubrics: Define sampling strategies, detailed annotation guidelines, and scoring rubrics. Measure inter‑annotator agreement and manage label noise to ensure benchmark integrity over time.

  • Deep Error & Trajectory Analysis: Conduct hands‑on failure analysis on agent runs, cluster root causes into actionable error categories, and clearly communicate findings to both engineering teams and stakeholders.

  • Drive Harness Scaffolding Iteration: Translate evaluation results directly into architectural improvements across agent scaffolding, prompt design, tool execution loops, context window compaction, retries, and safety guardrails.


What we are looking for:


  • Grounding in Statistics & Experimental Design: Demonstrated expertise in hypothesis testing, bootstrapping, power analysis, and statistical significance within probabilistic machine learning systems.

  • LLM‑as‑a‑Judge Validation: Proven experience building, auditing, and validating LLM judge setups against human annotations, including tracking agreement metrics and mitigating judge biases.

  • Evaluation Dataset Engineering: Strong track record designing evaluation datasets, crafting precise scoring rubrics, managing label noise, and measuring inter‑annotator agreement.

  • Evaluation of Non‑Deterministic Multi‑Step Systems: Deep comfort evaluating complex agent trajectories, tool execution sequences, latency, cost, and variance across repeated runs rather than relying solely on single‑shot benchmarks.

  • Metric Design & Integrity: Exceptional judgment in defining metrics tied to real‑world outcomes, with a keen eye for detecting benchmark gaming, target drift, or shortcut learning.

  • Rigor in Error Analysis: Ability to dissect complex execution logs, categorize failure modes, and synthesize clear, data‑driven recommendations.



  • Hands‑On LLM Integration: Practical proficiency with LLM APIs, prompt engineering, structured outputs, function/tool calling, and retrieval systems in Python.

  • Agent Architecture Awareness: Clear understanding of planning loops, tool orchestration, memory management, and multi‑agent coordination, including their trade‑offs and common failure modes.

  • Harness & Scaffolding Contribution: Ability to write clean, maintainable Python code to iterate on harness components such as guardrails, retry logic, state handling, and context management.


Nice to have:


  • Experience with agent and evaluation frameworks (e.g., LangChain, LangGraph, AutoGen, lm‑evaluation‑harvest, Promptfoo, or Ragas).

  • Exposure to running evaluation pipelines or agent workloads within secure, enclave, or Kubernetes environments.


Benefits:

Our benefits are designed to ensure our employees feel taken care of and are proud to be a part of the Valarian team. We are committed to consistently enhancing our benefit package, taking into account the overall well‑being and needs of our teammates. Here are the key benefits accessible to all employees at Valarian Technologies:



  • Equity – because you have the right to own what you’re building

  • A competitive salary – because we value your unique skills

  • Employer pension contributions – because you deserve a secure future

  • Private health insurance - because your health is important to us

  • Hybrid work setup – because everyone has different needs

  • Rewarding company retreats and meetups that respect your work/life balance – because we love getting to know each other!


Life at Valarian

Our culture is built on inclusivity, compassion and flexibility – we want everyone to be empowered to achieve their goals at Valarian.


The work we do is vital, but so are the connections that make it happen. We thrive on the shared energy, spontaneous conversations, and mutual trust built when we spend time together. We operate on a hybrid model, gathering in our London office 3 days a week to support one another and collaborate.


We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.


Valarian Technologies Limited is an equal opportunity employer and welcomes applications from individuals regardless of race, colour, religion, sex, sexual orientation, gender, identity or expression, national origin, age, disability, genetic information, marital status, veteran, amnesty, or any other legally protected characteristic.


We are committed to ensuring a fair and inclusive recruitment process and providing employment opportunities to all applicants. Decision recruitment, hiring, and employment are based solely on qualifications, skills, and experience relevant to the job requirements.


We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Engineering AI Engineer (Harness) London — Full time Apply →
Engineering AI Engineer (Harness) London — Full time Apply →

Valarian • Greater London

Hybrid
GBP 95,000 - 140,000
Equity
Competitive salary
Employer pension
+3
AI Engineer (Harness)
AI Engineer (Harness)

Valarian Technologies • Greater London

Hybrid
GBP 75,000 - 110,000
Equity
Competitive salary
Employer pension contributions
+3
Platform Engineer
Platform Engineer

Valarian • Greater London

Hybrid
GBP 90,000 - 125,000
Equity
Competitive salary
Employer pension contributions
+3
Senior Software Engineer
Senior Software Engineer

Valarian Technologies • Greater London

Hybrid
GBP 90,000 - 160,000
Equity
Competitive salary
Employer pension contributions
+3
FinOps Manager
FinOps Manager

Valarian • Greater London

Hybrid
GBP 70,000 - 110,000
Equity
Competitive salary
Employer pension contributions
+5
Software Engineer
Software Engineer

Valarian Technologies • Greater London

Hybrid
GBP 80,000 - 120,000
Equity
Competitive salary
Employer pension contributions
+3
Senior Technical Product Manager
Senior Technical Product Manager

Valarian • Greater London

On-site
GBP 90,000 - 120,000
Equity
Competitive salary
Employer pension contributions
+3
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Valarian • Greater London

Hybrid
GBP 95,000 - 150,000
Equity
Salary
Employer pension contributions
+3
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Valarian Technologies • Greater London

Hybrid
GBP 90,000 - 130,000
Equity
Competitive salary
Employer pension contributions
+3
Release Engineer
Release Engineer

Valarian Technologies Limited • Greater London

Hybrid
GBP 95,000 - 135,000
Equity
Private health insurance
Hybrid work setup
+2