AI QA Evaluation Engineer: Scale & Improve AI Answers

Appnovation Technologies

Toronto

On-site

CAD 90,000 - 130,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Appnovation Technologies is seeking a QA / AI Evaluation Engineer to join a fast-paced, scale-driven team. You will run evaluations at scale, measure factual grounding and accuracy lift, and build metrics frameworks that show improvement.

The role emphasizes automation, CI/CD integration, and collaboration with engineering to reproduce and verify fixes, while maintaining quality gates and test plans. We value experience with ML systems, data science techniques, and comfort across multiple LLM

Qualifications

  • Bachelor's degree in a technical field or equivalent experience.
  • 4+ years in QA / test engineering, with exposure to data/ML systems.
  • Strong Python and data-science techniques for measuring factual grounding and answer quality.
  • Experience with LLM evaluation frameworks and statistical analysis.
  • Ability to build automated, large-scale eval harnesses plus lighter human-in-the-loop A/B tests.
  • Comfort working across multiple LLM providers' outputs.
  • Test automation frameworks and scripting.
  • Detail-oriented, with strong analytical and communication skills.

Responsibilities

  • Run evals at scale across large question sets, from small human-UAT batches up to hundreds of thousands or millions of automated evaluations.
  • Statistically measure factual grounding and accuracy lift (before/after), not just human A/B testing.
  • Build a metrics framework showing quality improvement (e.g., "answer is X% supported by source content / Y% better").
  • Design load and quality tests as the corpus scales.
  • Define and maintain test plans, test cases, and quality gates.
  • Automate regression and evaluation suites; integrate them into CI/CD.
  • Report quality metrics clearly to technical and non-technical stakeholders.
  • Collaborate with engineering to reproduce, triage, and verify fixes.
  • Continuously improve QA processes and coverage.

Skills

Python
Data science
LLM eval frameworks
QA automation
Statistical analysis

Education

Bachelor's Degree in a technical field
Experience with data/ML systems

Tools

CI/CD
Python tooling
LLM providers' outputs

Job description

Appnovation Technologies is seeking a QA / AI Evaluation Engineer to join a fast-paced, scale-driven team. You will run evaluations at scale, measure factual grounding and accuracy lift, and build metrics frameworks that show improvement.

The role emphasizes automation, CI/CD integration, and collaboration with engineering to reproduce and verify fixes, while maintaining quality gates and test plans. We value experience with ML systems, data science techniques, and comfort across multiple LLM

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI QA Evaluation Engineer — Scale & Metrics Expert
AI QA Evaluation Engineer — Scale & Metrics Expert

Appnovation • Toronto

On-site
CAD 85,000 - 110,000
Diversity and Inclusion program
AI Evaluation Engineer — Scale QA & Metrics
AI Evaluation Engineer — Scale QA & Metrics

Socket.dev • Toronto

On-site
CAD 90,000 - 130,000
Senior AI QA & Evaluation Engineer
Senior AI QA & Evaluation Engineer

Socket.dev • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation Technologies • Toronto

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Socket.dev • Toronto

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Appnovation • Toronto

On-site
CAD 85,000 - 110,000
Diversity and Inclusion program
AI Quality Engineer
AI Quality Engineer

Centraprise • Vancouver

On-site
CAD 80,000 - 100,000
AI Evaluation Engineer (QA)
AI Evaluation Engineer (QA)

Socket.dev • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Lead QA Engineer - AI-Driven Automation & Performance
Lead QA Engineer - AI-Driven Automation & Performance

enableIT • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer - Agentic AI Quality & Metrics
AI Evaluation Engineer - Agentic AI Quality & Metrics

Doist • Kitchener

On-site
CAD 104,000 - 120,000
Ontario base salary