AI Evaluation Engineer

United States Digital Space LLC

Kitchener

On-site

CAD 90,000 - 130,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC seeks an AI Evaluation Engineer to join our AI Evaluation team in Canada. You will own evaluation coverage for the company’s Agentic AI systems alongside the evaluation lead, focusing on LLM-judge metrics, scenario and benchmark dataset curation, and error analysis to support release-readiness for voice and chat solutions.

This role reports to the AI Evaluation manager and may be based in our Vancouver office.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Software Engineering, Computational Linguistics, or a related field.
  • 3+ years of experience in QA, test engineering, model evaluation, or applied ML quality for AI-driven products.
  • Experience designing structured test strategies across manual and automated workflows.
  • Comfort working with complex AI systems such as speech, NLP, LLM, or agentic products.
  • Experience working with evaluation datasets, gold sets, adversarial test sets, or benchmark creation for AI systems.
  • Strong analytical skills for investigating failures, comparing outputs, and identifying actionable quality patterns.
  • Experience collaborating with cross-functional technical teams and communicating clearly through documentation and reporting.

Responsibilities

  • Design and execute validation strategies for agentic, NLP, and speech workflows across staging, beta, and release candidates.
  • Build, run, and improve regression evaluations, A/B comparisons, and red teaming analyses to determine readiness of product and model changes.
  • Co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions.
  • Create, configure, and monitor data annotation jobs to feed evaluation datasets on schedule.
  • Develop and maintain QA tooling, notebooks, and pipeline components for scalable evaluations.
  • Investigate bugs, triage issues, and decide on engineering escalations or follow-up analysis.
  • Collaborate with cross-functional teams, including applied science, engineering, and Product QA.

Skills

QA
Test strategy
AI evaluation
NLP
Data annotation
Cross-functional collaboration
Analytical thinking

Education

Bachelor's or Master's in CS/SE/Computational Linguistics

Job description

About the company

the company is the AI platform for customer experience, built to resolve customer problems in real time across voice and digital. Our AI agents learn from your best human agents and improve with every interaction, helping organizations understand their customers, deliver better experiences, increase operational efficiencies, and build a lasting competitive advantage.

Unlike legacy systems built to route and answer, or standalone agentic bot vendors built to deflect, the company was built to resolve. Our AI agents and human agents operate on a single platform with shared context, allowing Agentic AI to resolve issues, advance deals, and eliminate busywork through automation while seamlessly handing conversations to humans when needed, with full context preserved.

Market-leading brands, including Randstad, Motorola Solutions, Netflix, the San Diego Padres, the Colorado Rockies Baseball Club, and Cal Athletics, trust the company. the company is backed by Andreessen Horowitz, GV, ICONIQ Capital, and T-Mobile.

Being a Dialer

At the company, AI isn’t just a feature; it’s how our teams do their best work every day. We put powerful AI tools in every employee’s hands so they can move faster, think bigger, and achieve more.

We believe every conversation matters. And we’ve built the platform that turns those conversations into insight and action, for our customers and ourselves.

We look for people who are intensely curious and hold themselves to a high bar. Our ambition is significant, and achieving it requires a team that operates at the highest level. We seek individuals who embody our core traits: Scrappy, Curious, Optimistic, Persistent, and Empathetic.

Your role

As an AI Evaluation Engineer, you'll be an integral part of our AI Evaluation team, owning evaluation coverage for the company’s Agentic AI systems alongside our existing evaluation lead. A key focus will be co-owning LLM-judge metric development and calibration, scenario and benchmark dataset curation, and structured error analysis to support release-readiness decisions for our agentic voice and chat solutions.

This position reports to the manager of the AI Evaluation team and has the opportunity to be based in our Vancouver office.

What you’ll do
  • You will design and execute validation strategies for agentic, NLP, and speech workflows across staging, beta, and release candidates.
  • You will build, run, and improve regression evaluations, A/B comparisons, and red teaming analyses to determine whether product and model changes are ready to move forward.
  • You will co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions.
  • You will create, configure, and monitor data annotation jobs to keep evaluation and calibration datasets fed on schedule.
  • You will develop and maintain QA tooling, notebooks, and pipeline components that make recurring evaluations scalable and reusable across teams.
  • You will investigate bugs, triage issues, and decide whether problems should become engineering escalations, test set additions, or follow-up analysis.
  • You will collaborate with cross-functional teams, including applied science, engineering, and Product QA.
Skills you’ll bring
  • Bachelor's or Master's degree in Computer Science, Software Engineering, Computational Linguistics, or a related field.
  • 3+ years of experience in QA, test engineering, model evaluation, or applied ML quality for AI-driven products.
  • Experience designing structured test strategies across manual and automated workflows.
  • Comfort working with complex AI systems such as speech, NLP, LLM, or agentic products.
  • Experience working with evaluation datasets, gold sets, adversarial test sets, or benchmark creation for AI systems.
  • Strong analytical skills for investigating failures, comparing outputs, and identifying actionable quality patterns.
  • Experience collaborating with cross-functional technical teams and communicating clearly through documentation and reporting.

For exceptional talent based in Ontario, Canadathe target base salary range for this position is posted below. Our salary ranges are determined by role, level, and location. The range displayed on each job posting reflects the

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer
AI Evaluation Engineer

Dialpad Japan • Kitchener

On-site
CAD 104,000 - 120,000
Competitive salary
Comprehensive benefits
Training program
AI Evaluation Engineer
AI Evaluation Engineer

Dialpad Japan • Vancouver

On-site
CAD 115,000 - 133,000
Applied Scientist
Applied Scientist

United States Digital Space LLC • Kitchener

On-site
CAD 110,000 - 150,000
AI Evaluation Engineer - Agentic AI Quality & Metrics
AI Evaluation Engineer - Agentic AI Quality & Metrics

Doist • Kitchener

On-site
CAD 104,000 - 120,000
Ontario base salary
AI Evaluation Engineer - QA, Metrics & Benchmarking
AI Evaluation Engineer - QA, Metrics & Benchmarking

United States Digital Space LLC • Kitchener

On-site
CAD 90,000 - 130,000
AI Engineer
AI Engineer

Dialpad • Kitchener

Hybrid
CAD 145,000 - 173,000
Competitive salary
Comprehensive benefits
Training programs
+2
AI Engineer, Agentic Voice (TTS)
AI Engineer, Agentic Voice (TTS)

Dialpad Japan • Vancouver

Hybrid
CAD 161,000 - 192,000
Competitive salary
Comprehensive benefits
Career growth
Sr. SDET
Sr. SDET

Dialpad • Kitchener

On-site
CAD 135,000 - 158,000
Bonus
Equity
Comprehensive benefits package
AI Engineer, Agentic Voice (TTS)
AI Engineer, Agentic Voice (TTS)

Dialpad • Kitchener

Hybrid
CAD 145,000 - 173,000
Competitive salary
Comprehensive benefits
Growth opportunities
AI Engineer
AI Engineer

Dialpad • Vancouver

Hybrid
CAD 161,000 - 192,000
Comprehensive benefits
Competitive salary
Opportunities for growth