AI QA Evaluation Engineer — Scale & Metrics Expert

Appnovation

Toronto

On-site

CAD 85,000 - 110,000

Full time

31 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Diversity and Inclusion program

Job summary

Appnovation is seeking a QA / AI Evaluation Engineer to join a forward-leaning team. You will run large-scale evaluations, measure factual grounding and accuracy lift, and build a metrics framework to show quality improvements over time.

You will design tests, automate regressive suites, and collaborate with engineering to reproduce fixes. Strong Python, data-science skills, and experience with LLM evaluation are essential.

Qualifications

  • 4+ years in QA / test engineering with exposure to data/ML systems.
  • Strong Python and data-science techniques for measuring factual grounding and answer quality.
  • Experience with LLM evaluation frameworks and statistical analysis.
  • Ability to build automated, large-scale eval harnesses plus lighter human-in-the-loop A/B tests.
  • Comfort working across multiple LLM providers’ outputs.
  • Test automation frameworks and scripting.
  • Detail-oriented with strong analytical and communication skills.

Responsibilities

  • Run evals at scale across large question sets, from small human-UAT batches up to hundreds of thousands or millions of automated evaluations.
  • Statistically measure factual grounding and accuracy lift (before/after), not just human A/B testing.
  • Build a metrics framework showing quality improvement (e.g., “answer is X% supported by source content / Y% better”).
  • Design load and quality tests as the corpus scales.
  • Define and maintain test plans, test cases, and quality gates.
  • Automate regression and evaluation suites; integrate them into CI/CD.
  • Report quality metrics clearly to technical and non-technical stakeholders.
  • Collaborate with engineering to reproduce, triage, and verify fixes.
  • Continuously improve QA processes and coverage.

Skills

Python
Data science
LLM evaluation
Test automation
Scripting
CI/CD
Statistics
Analytical thinking

Education

Bachelor's Degree in a technical field

Tools

LLM evaluation frameworks
Python tooling
A/B testing
CI/CD pipelines

Job description

Appnovation is seeking a QA / AI Evaluation Engineer to join a forward-leaning team. You will run large-scale evaluations, measure factual grounding and accuracy lift, and build a metrics framework to show quality improvements over time.

You will design tests, automate regressive suites, and collaborate with engineering to reproduce fixes. Strong Python, data-science skills, and experience with LLM evaluation are essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI QA Evaluation Engineer: Scale & Improve AI Answers
AI QA Evaluation Engineer: Scale & Improve AI Answers

Appnovation Technologies • Toronto

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer — Scale QA & Metrics
AI Evaluation Engineer — Scale QA & Metrics

Socket.dev • Toronto

On-site
CAD 90,000 - 130,000
Senior AI QA & Evaluation Engineer
Senior AI QA & Evaluation Engineer

Socket.dev • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation Technologies • Toronto

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Socket.dev • Toronto

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation • Toronto

On-site
CAD 85,000 - 110,000
Diversity and Inclusion program
AI Evaluation Engineer - Agentic AI Quality & Metrics
AI Evaluation Engineer - Agentic AI Quality & Metrics

Doist • Kitchener

On-site
CAD 104,000 - 120,000
Ontario base salary
AI Quality Engineer
AI Quality Engineer

Centraprise • Vancouver

On-site
CAD 80,000 - 100,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Socket.dev • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer - QA, Metrics & Benchmarking
AI Evaluation Engineer - QA, Metrics & Benchmarking

United States Digital Space LLC • Kitchener

On-site
CAD 90,000 - 130,000