Senior AI Quality Engineer — LLM & Agentic Testing

Novara

Ontario (CA)

On-site

USD 70,000 - 92,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical
Dental
Vision
FSA
Paid holidays
Floating holidays
401k match
Life Insurance
Employee Assistance
Mental Health

Job summary

Novara is hiring a QA/Automation Engineer to own evaluation of AI-powered features and agent-based workflows. You will design frameworks for LLM outputs, build test harnesses, and integrate evals into CI/CD to safeguard production quality.

You will work with product, engineering, and QA teams to ensure robust testing of LLM-driven and agentic systems while expanding automation coverage across product flows. A hands-on IC role with high visibility in AI initiatives.

Qualifications

  • 6+ years in QA, SDET, or test automation, with real production automation shipping.
  • Hands-on experience testing LLM-based or agentic systems: building evals, working with LLM-as-judge patterns, prompt regression testing, or agent trajectory analysis.
  • Prior experience in a shift-left, embedded QA model.
  • Comfort with at least one modern automation stack (Playwright, Cypress, or similar) and a typed language, TypeScript preferred, Python fine.
  • Deep experience with test frameworks such as vitest, jest, or pytest, and comfort building custom test harnesses rather than only running off-the-shelf suites.
  • API-first testing mindset, including REST and Postman or equivalent.
  • Fluency in HTTP-level API testing, including recording proxies and observing service-to-service traffic.
  • Working knowledge of CI/CD pipelines, GitHub Actions a plus, and how to plug evals into them.
  • Familiarity with cloud secret managers and disciplined handling of sensitive test data in restore-from-prod environments.
  • Ability to reason clearly about probabilistic systems: variance, sample sizes, confidence, and when a flaky result is signal rather than noise.

Responsibilities

  • Design and maintain evaluation frameworks for LLM outputs and agentic workflows, including regression suites, golden datasets, and scoring rubrics.
  • Build test harnesses that catch hallucinations, tool-calling failures, prompt regressions, and unsafe or off-policy behavior before they reach production.
  • Define measurable quality criteria for agent reliability: task completion, factual grounding, latency, cost, and reasoning quality.
  • Integrate evaluation runs into CI/CD so model, prompt, and agent changes are gated the same way code changes are.
  • Partner with engineers on observability and tracing for agent runs, so failures are diagnosable rather than mysterious.
  • Contribute to conventional API and end-to-end automation where AI features sit inside larger product flows.
  • Help shape the team's shared playbook for testing AI features, and coach other QA engineers as agentic work spreads across scrum teams.

Skills

QA
SDET
Test automation
Playwright
TypeScript
Python
CI/CD
API testing
Observability
LangSmith

Tools

Postman
GitHub Actions
Datadog
Jira/Xray
AWS

Job description

Novara is hiring a QA/Automation Engineer to own evaluation of AI-powered features and agent-based workflows. You will design frameworks for LLM outputs, build test harnesses, and integrate evals into CI/CD to safeguard production quality.

You will work with product, engineering, and QA teams to ensure robust testing of LLM-driven and agentic systems while expanding automation coverage across product flows. A hands-on IC role with high visibility in AI initiatives.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

QA AI Automation Engineer - LLMs & Multi-Agent Quality
QA AI Automation Engineer - LLMs & Multi-Agent Quality

Dynasty Financial Partners, LLC • Saint Petersburg (FL)

On-site
USD 120,000 - 150,000
QA Engineer: AI & LLM Testing Automation
QA Engineer: AI & LLM Testing Automation

careers-ice • Atlanta (GA)

On-site
USD 90,000 - 150,000
QA AI Agent Developer — LLM-Driven Test Automation
QA AI Agent Developer — LLM-Driven Test Automation

Accord Technologies Inc • New Jersey

On-site
USD 90,000 - 120,000
Senior Agentic AI Engineer (LLM/Production)
Senior Agentic AI Engineer (LLM/Production)

Tryavalara • United States

On-site
USD 120,000 - 180,000
Paid time off
Parental leave
Health insurance
+1
Applied Scientist — AGI, LLM Quality & Evaluation
Applied Scientist — AGI, LLM Quality & Evaluation

Amazon • Bellevue (WA), Northern (KY)

Hybrid
USD 136,000 - 184,000
RSUs
Sign-on bonus
Health insurance
+2
Senior Agentic AI Engineer - LLM Systems & Production
Senior Agentic AI Engineer - LLM Systems & Production

Netomi • United States

Remote
USD 180,000 - 240,000
Lead QA Engineer for AI & LLM Workflows (Remote)
Lead QA Engineer for AI & LLM Workflows (Remote)

Vibehackers • Northern (KY)

Hybrid
USD 150,000 - 190,000
Generous annual bonus opportunity
401(k) with employer match
Medical Insurance
+2
LLM QA Engineer: AI Testing & Evaluation
LLM QA Engineer: AI Testing & Evaluation

Codefeast • United States

On-site
USD 90,000 - 140,000
QA / Automation Engineer Agentic AI
QA / Automation Engineer Agentic AI

Compunnel, Inc. • Atlanta (GA), Northern (KY)

On-site
USD 110,000 - 160,000
AI QA Engineer for LLM & GenAI Testing & Automation
AI QA Engineer for LLM & GenAI Testing & Automation

Compunnel, Inc. • New York (NY), Northern (KY)

Hybrid
USD 140,000 - 190,000