AI Evaluation Engineer — Scale QA & Metrics

Socket.dev

Toronto

On-site

CAD 90,000 - 130,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Appnovation Technologies in Toronto is seeking a QA / AI Evaluation Engineer to join a forward-leaning team that delivers high-quality AI-assisted solutions.

You will run evaluations at scale across large question sets, measure factual grounding and accuracy lift, and build the metrics framework that shows incremental improvement. You will design load tests and automate evaluation pipelines within CI/CD, collaborating with engineering and data science to reproduce and verify fixes.

Qualifications

  • Bachelor’s Degree in a technical field or equivalent experience.
  • 4+ years in QA / test engineering, with exposure to data/ML systems.
  • Strong Python and data-science techniques for measuring factual grounding and answer quality.
  • Experience with LLM evaluation frameworks and statistical analysis.
  • Ability to build automated, large-scale eval harnesses plus lighter human-in-the-loop A/B tests.
  • Comfort working across multiple LLM providers’ outputs.
  • Test automation frameworks and scripting.
  • Detail-oriented, with strong analytical and communication skills.

Responsibilities

  • Run evals at scale across large question sets, from small human-UAT batches up to hundreds of thousands or millions of automated evaluations.
  • Statistically measure factual grounding and accuracy lift (before/after), not just human A/B testing.
  • Build a metrics framework showing quality improvement (e.g., "answer is X% supported by source content / Y% better").
  • Design load and quality tests as the corpus scales.
  • Define and maintain test plans, test cases, and quality gates.
  • Automate regression and evaluation suites; integrate them into CI/CD.
  • Report quality metrics clearly to technical and non-technical stakeholders.
  • Collaborate with engineering to reproduce, triage, and verify fixes.
  • Continuously improve QA processes and coverage.

Skills

Python
Data science techniques
LLM evaluation frameworks
Statistical analysis
Test automation frameworks
Scripting
Communication skills

Education

Bachelor’s Degree in a technical field

Job description

Appnovation Technologies in Toronto is seeking a QA / AI Evaluation Engineer to join a forward-leaning team that delivers high-quality AI-assisted solutions.

You will run evaluations at scale across large question sets, measure factual grounding and accuracy lift, and build the metrics framework that shows incremental improvement. You will design load tests and automate evaluation pipelines within CI/CD, collaborating with engineering and data science to reproduce and verify fixes.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI QA Evaluation Engineer — Scale & Metrics Expert
AI QA Evaluation Engineer — Scale & Metrics Expert

Appnovation • Toronto

On-site
CAD 85,000 - 110,000
Diversity and Inclusion program
AI QA Evaluation Engineer: Scale & Improve AI Answers
AI QA Evaluation Engineer: Scale & Improve AI Answers

Appnovation Technologies • Toronto

On-site
CAD 90,000 - 130,000
Senior AI QA & Evaluation Engineer
Senior AI QA & Evaluation Engineer

Socket.dev • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer - QA, Metrics & Benchmarking
AI Evaluation Engineer - QA, Metrics & Benchmarking

United States Digital Space LLC • Kitchener

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation Technologies • Toronto

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer - Agentic AI Quality & Metrics
AI Evaluation Engineer - Agentic AI Quality & Metrics

Doist • Kitchener

On-site
CAD 104,000 - 120,000
Ontario base salary
AI Evaluation Engineer
AI Evaluation Engineer

United States Digital Space LLC • Kitchener

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Socket.dev • Toronto

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer — Agentic AI Quality & Benchmarks
AI Evaluation Engineer — Agentic AI Quality & Benchmarks

Dialpad • Vancouver

Hybrid
CAD 116,000 - 133,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation • Toronto

On-site
CAD 85,000 - 110,000
Diversity and Inclusion program