QA Engineer (AI Systems)

Engg

Toronto

On-site

CAD 100,000 - 150,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Nexxa.AI in Toronto seeks a Lead/Senior/Staff QA Engineer to own quality for AI agent systems that plan, call tools, and act in industrial environments. You will design evaluation frameworks for non-deterministic systems, build golden datasets, and stress-test agent behavior across planning, tool use, and recovery.

You will collaborate with ML and backend teams, develop robust testing pipelines, and mentor others on probabilistic system testing, ensuring observability and strong quality bars

Qualifications

  • 5+ years in QA/SDET roles owning test strategy for complex systems.
  • Hands-on experience testing LLM-based products, chatbots, or AI agents.
  • Practical experience with eval frameworks or building your own.
  • Strong scripting ability (Python) for test automation and data pipelines.
  • Understanding of LLM agents: prompting, tool calls, context management, and orchestration.

Responsibilities

  • Design and build evaluation harnesses and regression suites for LLM-based agents, covering reasoning quality, tool-call correctness, task completion, and multi-turn coherence.
  • Develop golden datasets and labeled test sets, including edge cases and adversarial prompts.
  • Define and track quality metrics beyond accuracy—groundedness, hallucination rate, task success, latency/cost tradeoffs, safety violations.
  • Build automated pipelines to run evals on every model, prompt, or tool integration change, and integrate into CI/CD.
  • Conduct structured red-teaming and adversarial testing with security teams.
  • Test agent behavior across the full action loop—planning, tool execution, error recovery, and final output.
  • Triage failures and translate eval failures into actionable bug reports.
  • Establish quality bars and sign-off criteria for new agent capabilities before customers see them.
  • Mentor other engineers on testing probabilistic, LLM-driven systems.
  • Advocate for testability and observability from day one.

Skills

QA Ownership
LLM Testing
Python Scripting
Eval Frameworks
Prompting
Ambiguity Handling
Written Communication

Tools

promptfoo
DeepEval
RAGAS
LangSmith

Job description

ROLE OVERVIEW

We're looking for a Lead / Senior / Staff QA Engineer to own quality for Nexxa's AI agent systems — products that plan, call tools, and take multi-step actions autonomously in industrial environments. This isn't traditional UI testing: you'll be designing evaluation frameworks for non-deterministic, tool-using systems, building golden datasets, catching regressions in reasoning quality, and stress-testing agent behavior under adversarial and real-world edge-case conditions. You'll work closely with ML engineers, backend engineers, and Forward Deployed Engineers to define what "good" looks like for an agent operating in high-stakes industrial settings, then build the infrastructure and processes to measure it continuously.

KEY RESPONSIBILITIES
  • Design and build evaluation harnesses and regression suites for LLM-based agents, covering reasoning quality, tool-call correctness, task completion, and multi-turn coherence.
  • Develop golden datasets and labeled test sets, including edge cases, ambiguous inputs, and adversarial prompts specific to industrial and operational contexts.
  • Define and track quality metrics beyond simple accuracy — groundedness, hallucination rate, task success rate, latency/cost tradeoffs, and safety violations.
  • Build automated pipelines that run evals on every model, prompt, or tool-integration change, and integrate them into CI/CD.
  • Conduct structured red-teaming and adversarial testing (prompt injection, jailbreaks, tool misuse, unsafe actions) in partnership with security teams.
  • Test agent behavior across the full action loop — planning, tool selection, tool execution, error recovery, and final output — not just the final response.
  • Investigate and triage failures where the root cause could be the model, the prompt, the tool/API, or the orchestration logic.
  • Partner with ML and backend engineers to translate eval failures into actionable, reproducible bug reports.
  • Establish quality bars and sign-off criteria for new agent capabilities before they reach customer environments.
  • Mentor other engineers on testing strategies specific to probabilistic, LLM-driven systems.
  • Advocate for testability and observability in agent architecture from day one.
QUALIFICATIONS
  • 5+ years in QA/SDET roles, with demonstrated ownership of test strategy for complex systems.
  • Hands-on experience testing LLM-based products, chatbots, or AI agents — you understand why traditional deterministic test assertions break down for generative systems.
  • Practical experience with eval frameworks or tooling (e.g., promptfoo, DeepEval, RAGAS, LangSmith) or a track record of building your own.
  • Strong scripting/programming ability (Python preferred) to build test automation, data pipelines, and eval tooling.
  • Understanding of how LLM agents work: prompting, tool/function calling, context management, RAG, memory, and orchestration frameworks.
  • Experience designing test data and labeled datasets, including sourcing, sampling, and managing dataset drift over time.
  • Familiarity with LLM-specific failure modes: hallucination, prompt injection, context poisoning, tool misuse, goal drift, and non-determinism.
  • Comfortable operating in ambiguity — defining what "correct" means for a task when there's no single right answer.
  • Strong written communication skills for turning fuzzy quality signals into clear, actionable findings for engineering and product stakeholders.
PREFERRED
  • Experience with human-in-the-loop evaluation workflows (labeling pipelines, inter-rater reliability, rubric design).
  • Background in ML/data science sufficient to read model evals and statistical significance.
  • Experience red-teaming or doing adversarial/security testing on ML systems.
  • Familiarity with observability/tracing tools for LLM applications (e.g., LangSmith, Arize, Langfuse, Weights & Biases).
  • Experience testing AI systems in industrial, IoT, or operational technology (OT) environments.
  • Prior experience setting up eval infrastructure from scratch at a startup or fast-moving team.
What We're Looking For
  • A QA engineer who wants to define what quality means for autonomous, real-world AI systems.
  • Someone who can build rigorous evaluation infrastructure for problems that don't have a single right answer.
  • A systems thinker who enjoys turning ambiguous agent behavior into measurable, trustworthy signals.
  • A strong collaborator who partners well with ML engineers, backend engineers, and Forward Deployed teams.
WHY JOIN NEXXA.AI http://Nexxa.AI?
  • Innovative Environment: Play a critical role in transforming heavy industries through groundbreaking AI and automation technologies.
  • Collaborative Culture: Be part of a team that values innovation, discipline, and continuous improvement.
  • Professional Growth: Benefit from significant opportunities for career development and advancement.
  • Competitive Compensation: Enjoy a comprehensive salary and equity package reflective of your expertise and contributions.

If you're passionate about AI quality and eager to help define what trustworthy autonomous systems look like in heavy industry, we'd love to connect.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Backend AI Engineer
Backend AI Engineer

Engg • Toronto

On-site
CAD 110,000 - 150,000
Backend AI Engineer
Backend AI Engineer

Nexxa.ai • Canada

On-site
CAD 120,000 - 180,000
Senior QA Engineer — AI Agent Systems & Autonomy
Senior QA Engineer — AI Agent Systems & Autonomy

Engg • Toronto

On-site
CAD 100,000 - 150,000
Staff DevOps Engineer
Staff DevOps Engineer

Engg • Toronto

On-site
CAD 140,000 - 210,000
Competitive compensation package
Equity or stock options
Principal AI Quality Engineer
Principal AI Quality Engineer

Worky • Canada

Remote
CAD 120,000 - 170,000
Stock options
Remote work
Flexible hours
+1
Security & Infrastructure Engineer
Security & Infrastructure Engineer

Nexxa.ai • Canada

On-site
CAD 120,000 - 180,000
Competitive salary
Equity package
Professional growth opportunities
Staff DevOps Engineer
Staff DevOps Engineer

Nexxa.ai • Canada

On-site
CAD 120,000 - 180,000
Agentic AI Optimization Developer
Agentic AI Optimization Developer

Equifax, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Security & Infrastructure Engineer
Security & Infrastructure Engineer

Engg • Toronto

On-site
CAD 120,000 - 190,000
AI Engineer
AI Engineer

Valsoft Corporation • Canada

On-site
CAD 100,000 - 140,000