AI Quality Engineer

Centraprise

Vancouver

On-site

CAD 80,000 - 100,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Centraprise is seeking an experienced QA Engineer based in Vancouver, Canada. You will be responsible for building and maintaining automated tests for AI agent workflows and ensuring quality metrics are met. Strong Python skills and experience testing backend services are essential.

The ideal candidate will have over 4 years in quality assurance or test automation, along with experience using frameworks like pytest and Playwright. Join us in enhancing the quality of cutting-edge AI technology!

Qualifications

  • 4+ years experience in QA engineering or similar roles.
  • Strong Python experience with production-quality tests.
  • Ability to work independently and identify root causes.

Responsibilities

  • Build and maintain automated tests for AI workflows.
  • Partner with engineering to define quality metrics.
  • Test on-device agent behavior including anomaly detection.

Skills

QA engineering
Python
test automation
experience with modern test frameworks
experience testing APIs
debugging application code

Tools

pytest
unittest
Playwright

Job description

Key Responsibilities
  • Build and maintain automated tests for AI agent workflows, APIs, tools, telemetry backed analysis, remediation flows, and ticketing behavior.
  • Design evaluation suites for LLM and agentic behavior, including expected-answer checks, rubric-based grading, regression datasets, tool-call validation, and safety/approval checks.
  • Use or help implement evaluation frameworks such as Pydantic Evals / Pydantic AI, Strands Evals, LangSmith, DeepEval, Ragas, promptfoo, or similar tools.
  • Validate multi-turn support scenarios, clarification flows, knowledge retrieval, script/remediation recommendations, escalation paths, and failure handling.
  • Test on-device agent behavior where needed, including Windows service/tray behavior, telemetry collection, anomaly detection, local remediation handoff, logs, and resource impact.
  • Debug quality issues directly by reading logs, tracing requests, reproducing failures, and making small code/test changes without heavy engineering hand-holding.
  • Partner with engineering and product to define release gates, quality metrics, evaluation rubrics, and confidence thresholds for pilot readiness.
  • Contribute to CI quality checks, test fixtures, mocked integrations, regression suites, and test data management.
  • Identify risks in AI behavior, including hallucinated diagnosis, unsafe remediation suggestions, missing consent, weak ticket summaries, brittle tool use, and poor escalation behavior.
Responsibilities
  • 4+ years of experience in QA engineering, SDET, test automation, software engineering, or similar hands‑on quality roles.
  • Strong Python experience, including writing production‑quality tests and debugging application code.
  • Experience testing backend services, APIs, async workflows, integrations, logs, and distributed systems.
  • Demonstrated ability to work independently in a codebase, identify root causes, and make targeted fixes or test improvements.
  • Experience with modern test frameworks such as pytest, unittest, Playwright, xUnit, or similar.
  • Comfortable testing ambiguous AI/non‑deterministic behavior using datasets, rubrics, assertions, metrics, and regression baselines.
  • Ability to distinguish product bugs, prompt/model behavior issues, data issues, integration failures, and test harness problems.
  • Strong written communication for documenting repro steps, risks, test plans, and release readiness.
Qualifications
  • Direct experience with Pydantic, Pydantic AI, or Pydantic Evals.
  • Experience testing LLM applications, agentic systems, RAG, tool calling, MCP servers, or AI assistants.
  • Experience building eval datasets, LLM‑as‑judge flows, deterministic evaluators, and regression dashboards.
  • Windows endpoint, device telemetry, PowerShell, .NET, MQTT, or on‑device agent testing experience.
  • Experience with IT support workflows, ServiceNow, ticketing, endpoint management, or enterprise support tooling.
  • Ability to contribute small application changes, not only test changes.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Test Automation Architect
Lead AI Test Automation Architect

High Tech Genesis Inc. • Toronto

On-site
CAD 120,000 - 180,000
Applied AI Developer (Agent Evaluation)
Applied AI Developer (Agent Evaluation)

Autodesk • Toronto

On-site
CAD 110,000 - 160,000
AI / ML QA Engineer
AI / ML QA Engineer

Infotek Consulting Inc. • Toronto

On-site
CAD 85,000 - 110,000
Lead AI Test Automation Architect
Lead AI Test Automation Architect

ITMC Systems, Inc • Toronto

On-site
CAD 150,000 - 190,000
Quality Assurance Automation Engineer
Quality Assurance Automation Engineer

Tata Consultancy Services • Montreal (administrative region)

On-site
CAD 90,000 - 120,000
Junior Quality Engineering Developer
Junior Quality Engineering Developer

Resonaite • Toronto

On-site
CAD 52,000 - 76,000
QE Automation Engineer – Playwright, Selenium, Java, REST Assured, GenAI Testing
QE Automation Engineer – Playwright, Selenium, Java, REST Assured, GenAI Testing

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision! • Toronto

On-site
CAD 80,000 - 100,000
Copy of AI Engineer
Copy of AI Engineer

PureFacts Financial Solutions • Toronto

On-site
CAD 90,000 - 140,000
Associate Director, Quality Engineering (12 Months contract)
Associate Director, Quality Engineering (12 Months contract)

EQ Bank • Toronto

On-site
CAD 110,000 - 150,000
AI Quality Engineer
AI Quality Engineer

Rootly • Toronto

Hybrid
CAD 80,000 - 100,000
Competitive compensation
Medical, dental, and vision coverage
3 weeks vacation and unlimited sick days
+1