AI Evaluation QA Engineer - Scale ML Testing & Metrics

Appnovation

Greater London

On-site

GBP 60,000 - 90,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Appnovation in London is seeking a QA / AI Evaluation Engineer to join a forward-leaning team that validates platform improvements through large-scale evaluations. You will measure factual grounding and accuracy lift while building reusable metrics to show progress over time.

The role requires 4+ years in QA for data/ML systems, strong Python skills, and experience with LLM evaluation frameworks. Collaboration with engineers to reproduce fixes is essential.

Qualifications

  • 4+ years in QA / test engineering with exposure to data/ML systems.
  • Strong Python and data-science techniques for measuring factual grounding and answer quality.
  • Experience with LLM evaluation frameworks and statistical analysis.
  • Ability to build automated, large-scale eval harnesses plus lighter human-in-the-loop A/B tests.
  • Comfort working across multiple LLM providers’ outputs.
  • Test automation frameworks and scripting.
  • Detail-oriented, with strong analytical and communication skills.

Responsibilities

  • Run evals at scale across large question sets, from small human-UAT batches up to hundreds of thousands or millions of automated evaluations.
  • Statistically measure factual grounding and accuracy lift (before/after), not just human A/B testing.
  • Build a metrics framework showing quality improvement (e.g., “answer is X% supported by source content / Y% better”).
  • Design load and quality tests as the corpus scales.
  • Define and maintain test plans, test cases, and quality gates.
  • Automate regression and evaluation suites; integrate them into CI/CD.
  • Report quality metrics clearly to technical and non-technical stakeholders.
  • Collaborate with engineering to reproduce, triage, and verify fixes.
  • Continuously improve QA processes and coverage.

Skills

Python
Data science techniques
LLM evaluation frameworks
Statistical analysis
Scripting
CI/CD
Test automation
Communication skills

Education

Bachelor’s Degree in a technical field

Tools

Python tooling
CI/CD pipelines

Job description

Appnovation in London is seeking a QA / AI Evaluation Engineer to join a forward-leaning team that validates platform improvements through large-scale evaluations. You will measure factual grounding and accuracy lift while building reusable metrics to show progress over time.

The role requires 4+ years in QA for data/ML systems, strong Python skills, and experience with LLM evaluation frameworks. Collaboration with engineers to reproduce fixes is essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI QA Evaluation Engineer: Scale Metrics for ML
AI QA Evaluation Engineer: Scale Metrics for ML

Appnovation Technologies • Greater London

On-site
GBP 60,000 - 90,000
AI Evaluation QA Engineer: Scale, Metrics & ML Quality
AI Evaluation QA Engineer: Scale, Metrics & ML Quality

Socket.dev • Greater London

On-site
GBP 55,000 - 75,000
Senior ML QA Engineer: Benchmark & Validate AI Systems
Senior ML QA Engineer: Benchmark & Validate AI Systems

EngineersOfAI • Greater London

On-site
GBP 60,000 - 80,000
Unlimited annual leave
Up to 5% matched pension
Phantom equity
+2
Senior ML QA Engineer - Build Reliable AI Stack (Flexible)
Senior ML QA Engineer - Build Reliable AI Stack (Flexible)

EngineersOfAI • Greater London

On-site
GBP 70,000 - 90,000
Unlimited annual leave
Up to 5% matched pension
Phantom equity
+4
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation Technologies • Greater London

On-site
GBP 60,000 - 90,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Socket.dev • Greater London

On-site
GBP 55,000 - 75,000
Senior ML QA Engineer - Build Reliable AI Stack (Flexible)
Senior ML QA Engineer - Build Reliable AI Stack (Flexible)

EngineersOfAI • Cambridge

Hybrid
GBP 60,000 - 80,000
Unlimited annual leave
Up to 5% matched pension
Phantom equity
+6
Senior ML QA Engineer: Benchmark & Validate AI Systems
Senior ML QA Engineer: Benchmark & Validate AI Systems

EngineersOfAI • Bristol

On-site
GBP 60,000 - 80,000
Unlimited annual leave
Up to 5% matched pension
Health cash plan
+2
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation • Greater London

On-site
GBP 60,000 - 90,000
Senior ML QA Engineer: Benchmark & Validate AI Systems
Senior ML QA Engineer: Benchmark & Validate AI Systems

EngineersOfAI • Cambridge

On-site
GBP 50,000 - 70,000
Unlimited annual leave
Up to 5% matched pension
Phantom equity
+4