AI Evaluation Engineer (QA)

Appnovation

Greater London

On-site

GBP 60,000 - 90,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Appnovation in London is seeking a QA / AI Evaluation Engineer to join a forward-leaning team that validates platform improvements through large-scale evaluations. You will measure factual grounding and accuracy lift while building reusable metrics to show progress over time.

The role requires 4+ years in QA for data/ML systems, strong Python skills, and experience with LLM evaluation frameworks. Collaboration with engineers to reproduce fixes is essential.

Qualifications

  • 4+ years in QA / test engineering with exposure to data/ML systems.
  • Strong Python and data-science techniques for measuring factual grounding and answer quality.
  • Experience with LLM evaluation frameworks and statistical analysis.
  • Ability to build automated, large-scale eval harnesses plus lighter human-in-the-loop A/B tests.
  • Comfort working across multiple LLM providers’ outputs.
  • Test automation frameworks and scripting.
  • Detail-oriented, with strong analytical and communication skills.

Responsibilities

  • Run evals at scale across large question sets, from small human-UAT batches up to hundreds of thousands or millions of automated evaluations.
  • Statistically measure factual grounding and accuracy lift (before/after), not just human A/B testing.
  • Build a metrics framework showing quality improvement (e.g., “answer is X% supported by source content / Y% better”).
  • Design load and quality tests as the corpus scales.
  • Define and maintain test plans, test cases, and quality gates.
  • Automate regression and evaluation suites; integrate them into CI/CD.
  • Report quality metrics clearly to technical and non-technical stakeholders.
  • Collaborate with engineering to reproduce, triage, and verify fixes.
  • Continuously improve QA processes and coverage.

Skills

Python
Data science techniques
LLM evaluation frameworks
Statistical analysis
Scripting
CI/CD
Test automation
Communication skills

Education

Bachelor’s Degree in a technical field

Tools

Python tooling
CI/CD pipelines

Job description

About Us

Appnovation is a global, full-service digital partner that combines Strategy, Experience & Design, Engineering and Managed Services. We build digital solutions that deliver real impact today and serve as foundations for future growth. Bold ambition. Practical action. Endless possibilities.

Technology is foundational to all of Appnovation’s offerings, from consulting to digital innovation, to digital product and service creation. The technology department is focused on delivering software solutions that enable rich consumer experiences, from mobile and web applications to advanced analytics and machine learning, to content and engagement management service enablement platforms.

Inherent throughout our tech capabilities is deep expertise in the Software Development Life Cycle, a drive for creativity, a passion for the craft, and collaboration with other disciplines - all foundational ingredients in successful digital experiences and client partnerships.

As a QA / AI Evaluation Engineer, you will join a highly motivated and experienced team in a forward-leaning role that proves the platform actually improves answer quality. You will run evaluations at scale — from small human-UAT batches up to millions of automated evals — statistically measure factual grounding and accuracy lift, and build the metrics framework that shows how much better our answers get over time. We are looking for people who can bring a strong, solution-focused mindset and contribute to quality standards, best practices and get things done.

Key Responsibilities
  • Run evals at scale across large question sets, from small human-UAT batches up to hundreds of thousands or millions of automated evaluations.
  • Statistically measure factual grounding and accuracy lift (before/after), not just human A/B testing.
  • Build a metrics framework showing quality improvement (e.g., “answer is X% supported by source content / Y% better”).
  • Design load and quality tests as the corpus scales.
  • Define and maintain test plans, test cases, and quality gates.
  • Automate regression and evaluation suites; integrate them into CI/CD.
  • Report quality metrics clearly to technical and non-technical stakeholders.
  • Collaborate with engineering to reproduce, triage, and verify fixes.
  • Continuously improve QA processes and coverage.
Qualifications
  • Bachelor’s Degree in a technical field or equivalent experience.
  • 4+ years in QA / test engineering, with exposure to data/ML systems.
  • Strong Python and data-science techniques for measuring factual grounding and answer quality.
  • Experience with LLM evaluation frameworks and statistical analysis.
  • Ability to build automated, large-scale eval harnesses plus lighter human-in-the-loop A/B tests.
  • Comfort working across multiple LLM providers’ outputs.
  • Test automation frameworks and scripting.
  • Detail-oriented, with strong analytical and communication skills.
Who You Are
  • You think about how to scale, automate and operate, not just how to build a solution to an immediate problem
  • You understand lean thinking
  • You set high standards for code quality, performance/scalability and security and seek continuous improvement
  • You have solid analytical, problem solving and decision-making skills
  • You have customer first mindset and a devotion to customer service
  • You engage and build positive internal and external client relationships, while managing multiple initiatives, often with competing priorities
  • You have strong self-initiative, passion, interpersonal, oral and written communication and collaboration skills with the ability to work, influence and make an impact in a cross-functional environment with all levels of the organization
  • You are responsive and thrive in a fast-paced diverse high-performance environment with rapidly changing business needs
  • You actively seek out things outside your comfort zone with the ability to rapidly learn and take advantage of new concepts, business models, and technologies
  • You have prior experience in consulting
  • Prior experience and connections in the Life Sciences industry is preferred

Thank you for your interest in a career with Appnovation Technologies! Please note that only those selected for an interview will be contacted.

At Appnovation, we recognize that diverse teams are the strongest teams. Diversity, Equity & Inclusion is not only something that we embrace - we celebrate it! We are proud to be an Equal Opportunity Employer and we encourage applicants from all backgrounds, lived experiences and industries to apply. Come join us at Appnovation, and learn more about how we stay true to our company values as we build better lives through better digital.

Accommodations are available upon request throughout the recruitment process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Socket.dev • Greater London

On-site
GBP 55,000 - 75,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation Technologies • Greater London

On-site
GBP 60,000 - 90,000
AI / Agent Engineer
AI / Agent Engineer

Appnovation Technologies • Greater London

On-site
GBP 90,000 - 130,000
AI / Agent Engineer
AI / Agent Engineer

Socket.dev • Greater London

On-site
GBP 90,000 - 150,000
AI Evaluation QA Engineer - Scale ML Testing & Metrics
AI Evaluation QA Engineer - Scale ML Testing & Metrics

Appnovation • Greater London

On-site
GBP 60,000 - 90,000
AI QA Evaluation Engineer: Scale Metrics for ML
AI QA Evaluation Engineer: Scale Metrics for ML

Appnovation Technologies • Greater London

On-site
GBP 60,000 - 90,000
AI Evaluation QA Engineer: Scale, Metrics & ML Quality
AI Evaluation QA Engineer: Scale, Metrics & ML Quality

Socket.dev • Greater London

On-site
GBP 55,000 - 75,000
AI Quality Lead
AI Quality Lead

Tripadvisor • Greater London

Hybrid
GBP 70,000 - 110,000
Competitive pay
Remote-friendly
Flexible schedule
+5
Quality Test Engineer III
Quality Test Engineer III

Elsevier Inc. • Greater London

On-site
GBP 60,000 - 90,000
Comprehensive Pension Plan
Generous vacation entitlement with sab
Maternity, Paternity, Adoption and ...
AI Quality Lead New London, England, United Kingdom
AI Quality Lead New London, England, United Kingdom

TripAdvisor LLC • Greater London

Hybrid
GBP 65,000 - 90,000
Competitive compensation packages
Remote-friendly collaboration
Flexible schedule
+5