Principal AI Quality Engineer

Worky

Canada

On-site

CAD 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock options
Remote work
Flexible hours
Mentorship program

Job summary

Nexaminds is seeking an AI Quality Engineer to lead validation and quality strategy for AI-powered systems, including LLMs, RAG applications, and agent-based workflows.

This role focuses on building enterprise-grade validation capabilities and partnering with engineering, AI/ML, and data governance teams to ensure AI solutions are accurate, safe, and production-ready.

Qualifications

  • 7+ years of experience in Quality Engineering or related technical discipline, including recent hands-on experience with AI/ML or LLM-based systems.
  • Experience designing validation frameworks for non-deterministic or probabilistic systems such as Large Language Models (LLMs) or Machine Learning applications.
  • Strong understanding of Retrieval-Augmented Generation (RAG), agentic workflows, and LLM evaluation concepts, including hallucination detection, groundedness, relevance, and consistency.
  • Hands-on experience with Large Language Models (LLMs) and prompt engineering, including prompt design, optimization, and regression validation.
  • Practical experience using Claude Code, Claude SDK, or similar AI-assisted development platforms and agentic coding tools.
  • Proficiency with CI/CD pipelines integration for AI validation.
  • Solid knowledge of data validation principles, including schema validation, business rule compliance, and data consistency.
  • Excellent cross-functional collaboration skills, with experience partnering across Quality Engineering, AI/ML Engineering, Software Engineering, and Data Governance teams.

Responsibilities

  • Define and lead the AI validation strategy for LLMs, RAG systems, and AI agents.
  • Build automated validation pipelines integrated into CI/CD.
  • Establish evaluation frameworks covering correctness, relevance, groundedness, consistency, and hallucination rate.
  • Design and maintain prompt regression testing to catch silent quality degradation after model or prompt changes.
  • Introduce semantic validation techniques that go beyond traditional QA (evaluating meaning, not just structure).
  • Lead production monitoring of AI output quality and define acceptable output ranges for non-deterministic systems.
  • Partner with engineering teams to validate AI-generated code, test cases, and AI-assisted decision workflows.
  • Introduce AI guardrails and confidence scoring to flag unsafe or low-confidence outputs.
  • Build validation checkpoints across multi-step agent pipelines.
  • Define and track quality metrics/KPIs (defect reduction, output consistency, prompt regression stability, adoption confidence).

Skills

Quality Engineering
AI/ML
LLMs
RAG concepts
Prompt engineering
CI/CD integration
Team collaboration
Evaluation pipelines

Tools

Node.js
MongoDB
Claude Code
Claude SDK
Playwright
RestSharp

Job description

At Nexaminds, we’re on a mission to redefine industries with AI. We’re passionate about the limitless potential of artificial intelligence to transform businesses, streamline processes, and drive growth.

Join us on our visionary journey. We’re leading the way in AI solutions, and we’re committed to innovation, collaboration, and ethical practices. Become a part of our team and shape the future powered by intelligent machines. If you’re driven by ambition, success, fun, and learning, Nexaminds is where you belong.

Nexaminds is looking for an AI Quality Engineer to lead the validation and quality strategy for AI-powered systems, including Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) applications, and agent-based workflows. The ideal candidate has deep experience in Quality Engineering, AI evaluation, and automated validation frameworks, with a strong understanding of how to measure, monitor, and improve the reliability of non-deterministic AI systems.This role focuses on building enterprise-grade AI validation capabilities, designing automated evaluation pipelines, implementing AI guardrails, and partnering closely with engineering, AI/ML, and data governance teams to ensure AI-powered solutions are accurate, safe, and production-ready.

Location: Canada (Remote)

Qualifications weare looking for:
  • 7+ years of experience in Quality Engineering, Software Development, or a related technical discipline, including recent hands-on experience with AI/ML or LLM-based systems.
  • Experience designing validation or evaluation frameworks for non-deterministic or probabilistic systems such as Large Language Models (LLMs) or Machine Learning applications.
  • Strong understanding of Retrieval-Augmented Generation (RAG), agentic workflows, and LLM evaluation concepts, including hallucination detection, groundedness, relevance, and consistency.
  • Hands-on experience with Large Language Models (LLMs) and prompt engineering, including prompt design, optimization, and regression validation.
  • Practical experience using Claude Code, Claude SDK, or similar AI-assisted development platforms and agentic coding tools.
  • Proficiency with Node.js for building automation tools, validation pipelines, and engineering utilities.
  • Experience working with MongoDB to store, manage, and analyze evaluation results, validation data, or AI output logs.
  • Experience integrating automated validation processes into CI/CD pipelines.
  • Experience designing semantic validation approaches that evaluate meaning, relevance, and contextual accuracy beyond traditional software testing.
  • Strong understanding of AI quality metrics, confidence scoring, output consistency, and production monitoring.
  • Solid knowledge of data validation principles, including schema validation, business rule compliance, and data consistency.
  • Excellent cross-functional collaboration skills, with experience partnering across Quality Engineering, AI/ML Engineering, Software Engineering, and Data Governance teams.
Nice to have:
  • Experience with prompt regression testing tools or frameworks.
  • Familiarity with AI governance, fairness, or bias-detection practices.
  • Experience with tools such as Playwright, RestSharp, or similar UI/API test automation.
  • Prior experience standing up a validation or evaluation function from scratch.
  • Exposure to confidence scoring or guardrail systems for production AI.
  • Experience designing or orchestrating multi-agent systems in production environments
Job duties:
  • Define and lead the AI validation strategy for LLMs, RAG systems, and AI agents.
  • Build automated validation pipelines integrated into CI/CD.
  • Establish evaluation frameworks covering correctness, relevance, groundedness, consistency, and hallucination rate.
  • Design and maintain prompt regression testing to catch silent quality degradation after model or prompt changes.
  • Introduce semantic validation techniques that go beyond traditional QA (evaluating meaning, not just structure).
  • Lead production monitoring of AI output quality and define acceptable output ranges for non-deterministic systems.
  • Partner with engineering teams to validate AI-generated code, test cases, and AI-assisted decision workflows.
  • Introduce AI guardrails and confidence scoring to flag unsafe or low-confidence outputs.
  • Build validation checkpoints across multi-step agent pipelines.
  • Define and track quality metrics/KPIs (defect reduction, output consistency, prompt regression stability, adoption confidence).
What you can expect from us

Here at Nexaminds, we’re not your typical workplace. We’re all about creating a friendly and trusting environment where you can thrive. Why does this matter? Well, trust and openness lead to better quality, innovation, commitment to getting the job done, efficiency, and cost-effectiveness.

  • Stock options
  • Remote work options
  • Flexible working hours
  • Benefits above the law
  • But it’s not just about the work; it’s about the people too. You’ll be collaborating with some seriously awesome IT pros.
  • You’ll have access to mentorship and tons of opportunities to learn and level up.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Quality Engineer
AI Quality Engineer

Rootly • Toronto

Hybrid
CAD 80,000 - 100,000
Competitive compensation
Medical, dental, and vision coverage
3 weeks vacation and unlimited sick days
+1
AI Engineer - Canada
AI Engineer - Canada

Pulsora • Canada

Remote
CAD 60,000 - 70,000
QA Lead — AI Systems & Models Testing
QA Lead — AI Systems & Models Testing

Jay Analytix • Montreal (administrative region)

Hybrid
CAD 100,000 - 130,000
AI Evaluation Engineer
AI Evaluation Engineer

United States Digital Space LLC • Kitchener

On-site
CAD 90,000 - 130,000
Manager, Software Engineering
Manager, Software Engineering

United States Digital Space LLC • Toronto

Hybrid
CAD 100,000 - 130,000
Uncapped vacation
Competitive benefits
Hybrid work environment
+1
AI Engineer, Sr 1
AI Engineer, Sr 1

Jobgether • Canada

On-site
CAD 100,000 - 160,000
Remote work flexibility
Comprehensive benefits package
Diverse, international team
+1
AI Developer
AI Developer

Genpact • Montreal (administrative region)

On-site
CAD 90,000 - 120,000
AI Learning & Enablement Lead
AI Learning & Enablement Lead

Jobgether • Canada

Remote
CAD 100,000 - 120,000
Discretionary incentive
Medical benefits
Fully remote within Canada
Senior Software Developer, Machine Learning, Applied AI
Senior Software Developer, Machine Learning, Applied AI

United States Digital Space LLC • Toronto

On-site
CAD 182,000 - 186,000
Equity
Benefits
Bonus target
Applied AI Engineer
Applied AI Engineer

Worky • Montreal (administrative region)

Hybrid
CAD 70,000 - 110,000