AI Evaluation Engineer - QA, Metrics & Benchmarking

United States Digital Space LLC

Kitchener

On-site

CAD 90,000 - 130,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC seeks an AI Evaluation Engineer to join our AI Evaluation team in Canada. You will own evaluation coverage for the company’s Agentic AI systems alongside the evaluation lead, focusing on LLM-judge metrics, scenario and benchmark dataset curation, and error analysis to support release-readiness for voice and chat solutions.

This role reports to the AI Evaluation manager and may be based in our Vancouver office.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Software Engineering, Computational Linguistics, or a related field.
  • 3+ years of experience in QA, test engineering, model evaluation, or applied ML quality for AI-driven products.
  • Experience designing structured test strategies across manual and automated workflows.
  • Comfort working with complex AI systems such as speech, NLP, LLM, or agentic products.
  • Experience working with evaluation datasets, gold sets, adversarial test sets, or benchmark creation for AI systems.
  • Strong analytical skills for investigating failures, comparing outputs, and identifying actionable quality patterns.
  • Experience collaborating with cross-functional technical teams and communicating clearly through documentation and reporting.

Responsibilities

  • Design and execute validation strategies for agentic, NLP, and speech workflows across staging, beta, and release candidates.
  • Build, run, and improve regression evaluations, A/B comparisons, and red teaming analyses to determine readiness of product and model changes.
  • Co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions.
  • Create, configure, and monitor data annotation jobs to feed evaluation datasets on schedule.
  • Develop and maintain QA tooling, notebooks, and pipeline components for scalable evaluations.
  • Investigate bugs, triage issues, and decide on engineering escalations or follow-up analysis.
  • Collaborate with cross-functional teams, including applied science, engineering, and Product QA.

Skills

QA
Test strategy
AI evaluation
NLP
Data annotation
Cross-functional collaboration
Analytical thinking

Education

Bachelor's or Master's in CS/SE/Computational Linguistics

Job description

United States Digital Space LLC seeks an AI Evaluation Engineer to join our AI Evaluation team in Canada. You will own evaluation coverage for the company’s Agentic AI systems alongside the evaluation lead, focusing on LLM-judge metrics, scenario and benchmark dataset curation, and error analysis to support release-readiness for voice and chat solutions.

This role reports to the AI Evaluation manager and may be based in our Vancouver office.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer - Agentic AI Quality & Metrics
AI Evaluation Engineer - Agentic AI Quality & Metrics

Doist • Kitchener

On-site
CAD 104,000 - 120,000
Ontario base salary
AI Evaluation Engineer
AI Evaluation Engineer

United States Digital Space LLC • Kitchener

On-site
CAD 90,000 - 130,000
AI Evaluation Engineer
AI Evaluation Engineer

Dialpad Japan • Vancouver

On-site
CAD 115,000 - 133,000
AI Evaluation Engineer
AI Evaluation Engineer

Dialpad Japan • Kitchener

On-site
CAD 104,000 - 120,000
Competitive salary
Comprehensive benefits
Training program
QA Lead — AI Systems & Models Testing
QA Lead — AI Systems & Models Testing

Jay Analytix • Montreal (administrative region)

Hybrid
CAD 100,000 - 130,000
AI Safety & Customer Care ML Engineer (Production Agents)
AI Safety & Customer Care ML Engineer (Production Agents)

United States Digital Space LLC • Toronto

Hybrid
CAD 119,000 - 149,000
Extended health and dental coverage
Mental health benefits
Family building benefits
+5
Principal AI Quality Engineer
Principal AI Quality Engineer

Worky • Canada

Remote
CAD 120,000 - 170,000
Stock options
Remote work
Flexible hours
+1
Senior Software Developer, Machine Learning, Applied AI
Senior Software Developer, Machine Learning, Applied AI

United States Digital Space LLC • Toronto

On-site
CAD 182,000 - 186,000
Equity
Benefits
Bonus target
AI-Powered CX Strategy Lead
AI-Powered CX Strategy Lead

United States Digital Space LLC • Vancouver

Hybrid
CAD 168,000 - 209,000
Hybrid work environment
Remote work up to 4 weeks/year
AI Enablement Engineer & Coach
AI Enablement Engineer & Coach

Electric-Mind • Toronto

On-site
CAD 100,000 - 150,000