Senior AI QA & Evaluation Engineer

Socket.dev

Montreal (administrative region)

On-site

CAD 90,000 - 130,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Appnovation is seeking a QA / AI Evaluation Engineer to lead large-scale evaluations of AI outputs and improve the accuracy and reliability of our platform. You will design metrics, run automated and human-in-the-loop tests, and collaborate with engineering to reproduce and fix issues.

This role emphasizes scalable QA practices, data-driven decision making, and a strong focus on model grounding and quality across teams.

Qualifications

  • : Bachelor’s degree or equivalent experience in a technical field.
  • -

Responsibilities

  • Run evals at scale across large question sets, from small human-UAT batches up to hundreds of thousands or millions of automated evaluations.
  • Statistically measure factual grounding and accuracy lift (before/after), not just human A/B testing.
  • Build a metrics framework showing quality improvement (e.g., "answer is X% supported by source content / Y% better").
  • Design load and quality tests as the corpus scales.
  • Define and maintain test plans, test cases, and quality gates.
  • Automate regression and evaluation suites; integrate them into CI/CD.
  • Report quality metrics clearly to technical and non-technical stakeholders.
  • Collaborate with engineering to reproduce, triage, and verify fixes.
  • Continuously improve QA processes and coverage.

Skills

Python
QA / testing
LLM evaluation
Statistical analysis
Automation
Communication

Education

Bachelor’s Degree in a technical field

Tools

LLM evaluation frameworks
Test automation frameworks
Scripting
CI/CD pipelines

Job description

Appnovation is seeking a QA / AI Evaluation Engineer to lead large-scale evaluations of AI outputs and improve the accuracy and reliability of our platform. You will design metrics, run automated and human-in-the-loop tests, and collaborate with engineering to reproduce and fix issues.

This role emphasizes scalable QA practices, data-driven decision making, and a strong focus on model grounding and quality across teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI QA Evaluation Engineer — Scale & Metrics Expert
AI QA Evaluation Engineer — Scale & Metrics Expert

Appnovation • Toronto

On-site
CAD 85,000 - 110,000
Diversity and Inclusion program
AI QA Evaluation Engineer: Scale & Improve AI Answers
AI QA Evaluation Engineer: Scale & Improve AI Answers

Appnovation Technologies • Toronto

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer — Scale QA & Metrics
AI Evaluation Engineer — Scale QA & Metrics

Socket.dev • Toronto

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation Technologies • Toronto

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Socket.dev • Toronto

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation • Toronto

On-site
CAD 85,000 - 110,000
Diversity and Inclusion program
AI Evaluation Engineer - Agentic AI Quality & Metrics
AI Evaluation Engineer - Agentic AI Quality & Metrics

Doist • Kitchener

On-site
CAD 104,000 - 120,000
Ontario base salary
AI / ML QA Engineer
AI / ML QA Engineer

Infotek Consulting Inc. • Toronto

On-site
CAD 85,000 - 110,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Socket.dev • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer - QA, Metrics & Benchmarking
AI Evaluation Engineer - QA, Metrics & Benchmarking

United States Digital Space LLC • Kitchener

On-site
CAD 90,000 - 130,000