AI QA Evaluation Engineer: Scale Metrics for ML

Appnovation Technologies

Greater London

On-site

GBP 60,000 - 90,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Appnovation Technologies is seeking a QA / AI Evaluation Engineer to join a forward-thinking data/ML focused team. You will run large-scale evals, measure factual grounding and accuracy lift, and build metrics frameworks to demonstrate quality improvements across AI outputs.

Applicants should have a strong background in QA/test engineering, Python, and experience with LLM evaluation frameworks, plus the ability to design automated evaluation harnesses and integrate them with CI/CD.

Qualifications

  • 4+ years in QA / test engineering, with exposure to data/ML systems.
  • Strong Python and data-science techniques for measuring factual grounding and answer quality.
  • Experience with LLM evaluation frameworks and statistical analysis.

Responsibilities

  • Run evals at scale across large question sets, from small batches to millions of evaluations.
  • Statistically measure factual grounding and accuracy lift (before/after).
  • Build a metrics framework showing quality improvements (X% supported by sources).
  • Design load and quality tests as corpus scales.
  • Define and maintain test plans, test cases, and quality gates.
  • Automate regression and evaluation suites; integrate into CI/CD.
  • Report quality metrics clearly to stakeholders.
  • Collaborate with engineering to reproduce and verify fixes.
  • Continuously improve QA processes and coverage.

Skills

QA testing
Python
LLM evaluation
Data/ML
Automation

Education

Bachelor’s Degree in a technical field

Tools

Python
CI/CD
LLM eval frameworks

Job description

Appnovation Technologies is seeking a QA / AI Evaluation Engineer to join a forward-thinking data/ML focused team. You will run large-scale evals, measure factual grounding and accuracy lift, and build metrics frameworks to demonstrate quality improvements across AI outputs.

Applicants should have a strong background in QA/test engineering, Python, and experience with LLM evaluation frameworks, plus the ability to design automated evaluation harnesses and integrate them with CI/CD.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

QA Engineer
QA Engineer

iFindTech Ltd • Greater London

On-site
GBP 60,000 - 90,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation Technologies • Greater London

On-site
GBP 60,000 - 90,000
AI Quality Engineer: Build Evaluation Frameworks & Automation
AI Quality Engineer: Build Evaluation Frameworks & Automation

iFindTech Ltd • Greater London

On-site
GBP 60,000 - 90,000
AI QA Engineer: Scale AI Quality & Evaluation (Hybrid)
AI QA Engineer: Scale AI Quality & Evaluation (Hybrid)

WeDo Technology Solutions Limited • Greater London

Hybrid
GBP 70,000 - 95,000
10% bonus
Employee shares/equity
Private healthcare
+3
AI Quality & Evaluation Engineer: Production Metrics
AI Quality & Evaluation Engineer: Production Metrics

ConnexAI • Manchester

On-site
GBP 50,000 - 70,000
AI QA Engineer
AI QA Engineer

WeDo Technology Solutions Limited • Greater London

Hybrid
GBP 70,000 - 95,000
10% bonus
Employee shares/equity
Private healthcare
+3
AI QA Engineer
AI QA Engineer

WeDoTech • Greater London

Hybrid
GBP 70,000 - 95,000
Employee shares/equity
Private healthcare
Life insurance
+2
Senior ML QA Engineer - Build Reliable AI Stack (Flexible)
Senior ML QA Engineer - Build Reliable AI Stack (Flexible)

EngineersOfAI • Cambridge

Hybrid
GBP 60,000 - 80,000
Unlimited annual leave
Up to 5% matched pension
Phantom equity
+6
Senior ML QA Engineer: Benchmark & Validate AI Systems
Senior ML QA Engineer: Benchmark & Validate AI Systems

EngineersOfAI • Bristol

On-site
GBP 60,000 - 80,000
Unlimited annual leave
Up to 5% matched pension
Health cash plan
+2
Senior ML QA Engineer - Build Reliable AI Stack (Flexible)
Senior ML QA Engineer - Build Reliable AI Stack (Flexible)

EngineersOfAI • Bristol

On-site
GBP 45,000 - 65,000
Unlimited annual leave
Up to 5% matched pension
Phantom equity
+5