Senior AI Quality Engineer

Novara

United States

On-site

USD 70,000 - 91,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical
Dental
Vision
Flexible Spending Accounts
PTO and holidays
401k with company match
Life Insurance
Employee Assistance Programs
No-cost Mental Health Benefits

Job summary

Novara is seeking a hands-on QA/SDET specialist to own AI evaluation for non-deterministic systems. You will design frameworks, build test harnesses, and gate AI changes in CI/CD, partnering with product and engineering to raise quality for LLM-driven workflows.

You will mentor teams, contribute to API and end-to-end automation, and shape the QA playbook as agentic features spread across scrum teams, with strong focus on reliability, grounding, and latency.

Qualifications

  • 6+ years in QA, SDET, or test automation, with real production automation shipping
  • Hands-on experience testing LLM-based or agentic systems: building evals, working with LLM-as-judge patterns, prompt regression testing, or agent trajectory analysis
  • Prior experience in a shift-left, embedded QA model
  • Comfort with at least one modern automation stack (Playwright, Cypress, or similar) and a typed language, TypeScript preferred, Python fine
  • Deep experience with test frameworks such as vitest, jest, or pytest, and comfort building custom test harnesses rather than only running off-the-shelf suites
  • API-first testing mindset, including REST and Postman or equivalent
  • Fluency in HTTP-level API testing, including recording proxies and observing service-to-service traffic
  • Working knowledge of CI/CD pipelines, GitHub Actions a plus, and how to plug evals into them
  • Familiarity with cloud secret managers and disciplined handling of sensitive test data in restore-from-prod environments
  • Ability to reason clearly about probabilistic systems: variance, sample sizes, confidence, and when a flaky result is signal rather than noise

Responsibilities

  • Design and maintain evaluation frameworks for LLM outputs and agentic workflows, including regression suites, golden datasets, and scoring rubrics
  • Build test harnesses that catch hallucinations, tool-calling failures, prompt regressions, and unsafe or off-policy behavior before they reach production
  • Define measurable quality criteria for agent reliability: task completion, factual grounding, latency, cost, and reasoning quality
  • Integrate evaluation runs into CI/CD so model, prompt, and agent changes are gated the same way code changes are
  • Partner with engineers on observability and tracing for agent runs, so failures are diagnosable rather than mysterious
  • Contribute to conventional API and end-to-end automation where AI features sit inside larger product flows
  • Help shape the team's shared playbook for testing AI features, and coach other QA engineers as agentic work spreads across scrum teams

Skills

QA automation
LLM testing
Agentic systems
Playwright
Cypress
Vitest/Jest/Pytest
TypeScript
REST API testing
CI/CD
GitHub Actions
Observability/tracing
Data handling in prod

Tools

Playwright
Cypress
Postman
Datadog
Jira/Xray
AWS

Job description

Position Description:
AI is core to where our product is heading, and the quality bar for it has to be as high as anything else we ship. This role owns that bar. You will define what "working" means for non-deterministic systems, build the tooling to prove it, and give our teams the confidence to move faster on AI-powered features and agent-based products.
You will partner closely with product, engineering, and the broader QA team to bring rigor to how we test LLM-driven and agentic workflows, while also contributing to traditional automation coverage where it matters.
This is a hands-on IC role reporting into the QA Manager, with high visibility into AI initiatives across the company.

Responsibilities:
  • Design and maintain evaluation frameworks for LLM outputs and agentic workflows, including regression suites, golden datasets, and scoring rubrics
  • Build test harnesses that catch hallucinations, tool-calling failures, prompt regressions, and unsafe or off-policy behavior before they reach production
  • Define measurable quality criteria for agent reliability: task completion, factual grounding, latency, cost, and reasoning quality
  • Integrate evaluation runs into CI/CD so model, prompt, and agent changes are gated the same way code changes are
  • Partner with engineers on observability and tracing for agent runs, so failures are diagnosable rather than mysterious
  • Contribute to conventional API and end-to-end automation where AI features sit inside larger product flows
  • Help shape the team's shared playbook for testing AI features, and coach other QA engineers as agentic work spreads across scrum teams
Knowledge, Experience, Requirements:
  • 6+ years in QA, SDET, or test automation, with real production automation shipping
  • Hands-on experience testing LLM-based or agentic systems: building evals, working with LLM-as-judge patterns, prompt regression testing, or agent trajectory analysis
  • Prior experience in a shift-left, embedded QA model
  • Comfort with at least one modern automation stack (Playwright, Cypress, or similar) and a typed language, TypeScript preferred, Python fine
  • Deep experience with test frameworks such as vitest, jest, or pytest, and comfort building custom test harnesses rather than only running off-the-shelf suites
  • API-first testing mindset, including REST and Postman or equivalent
  • Fluency in HTTP-level API testing, including recording proxies and observing service-to-service traffic
  • Working knowledge of CI/CD pipelines, GitHub Actions a plus, and how to plug evals into them
  • Familiarity with cloud secret managers and disciplined handling of sensitive test data in restore-from-prod environments
  • Ability to reason clearly about probabilistic systems: variance, sample sizes, confidence, and when a flaky result is signal rather than noise
Preferred Qualifications:
  • Experience with eval tooling such as Langfuse, Braintrust, LangSmith, Ragas, or DeepEval
  • Familiarity with RAG systems, vector stores, or tool-calling frameworks
  • Background in test data strategy for AI, including synthetic data generation
  • Exposure to Datadog or a similar observability platform
Tech Stack:

TypeScript, Playwright, vitest, Postman, GitHub Actions, Jira/Xray, Datadog, AWS, plus emerging AI evaluation tooling.

Compensation:

Annual Base Salary Range of 100k-130k CAD
Annual Bonus Opportunity of 10%

As a growing company, Novara values its employees by supporting them with a full benefits package including Medical, Dental, Vision, Flexible Spending Accounts, PTO, Paid and Floating Holidays, 401k with Company match and immediate vesting, Company-funded Life Insurance, Employee Assistance Programs, and No-cost Mental Health Benefits.

About Novara

Novara provides safety and operational risk management software that empowers organizations to identify and resolve issues before they become incidents. Through the Flex and Risk Management Center platforms, Novara helps organizations address operational risk proactively by unifying data, increasing workforce engagement, and proactively managing risk. Novara's combination of training, software, and tools puts people and safety first while protecting critical operations.

Novara, a Providence Equity portfolio company, provides safety and operational risk management software that empowers organizations to identify and resolve issues before they become incidents. Through the Flex and Risk Management Center platforms, Novara helps organizations address operational risk proactively by unifying data, increasing workforce engagement, and proactively managing risk. Novara's combination of training, software, and tools puts people and safety first while protecting critical operations.

Novara launched January 1 2026, as an independent company, a spinoff of the Flex and RMC software businesses formerly part of KPA.

Don't meet every job requirement? At Novara, we are dedicated to building a diverse, inclusive, and authentic workplace. Studies have shown that women and people of color are less likely to apply unless they meet every requirement. If you're excited about the role but your past experience doesn't align perfectly with every qualification, we still encourage you to apply! You might just be the right candidate for this or other roles.

Please note that we may use AI tools to assist in the initial screening of resumes to help identify qualified candidates more efficiently. All decisions are reviewed by a human recruiter, and no hiring determination is made solely by automated means.

Novara is committed to providing equal opportunity in all of our employment practices, including selection, hiring, promotion, transfer, and compensation, to all qualified applicants and employees without regard to race, religion, religious dress/grooming, color, ethnicity, sex (including sex stereotyping), sexual orientation, gender identity or gender expression, national origin, ancestry, citizenship status, creed, uniform service member status, military or veteran status, marital status, pregnancy, breast-feeding and/or pregnancy-related conditions, age, protected medical condition, leave status, physical or mental disability, genetic characteristics, or any other legally-protected status in accordance with the requirements of all federal, state and local laws. In compliance with federal law, all persons hired will be required to verify identity and eligibility to work in the United States and to complete the required employment eligibility verification document form upon hire.

If you need assistance or an accommodation due to a disability, you may contact us at hr@novara.com.

Please see our Candidate Privacy Notice Included Here

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer
Senior Software Engineer

Novara • United States

On-site
USD 130,000 - 150,000
Medical
Dental
Vision
+7
IT Support Specialist
IT Support Specialist

Novara • Westminster (CO)

Hybrid
USD 88,784,000 - 103,155,000
Medical benefits
Dental benefits
Vision benefits
+5
Account Executive
Account Executive

Novara • United States

On-site
USD 115,000 - 125,000
Medical
Dental
Vision
+4
Senior Corporate Counsel
Senior Corporate Counsel

Lever, Inc. • United States

Remote
USD 160,000 - 180,000
Medical
Dental
Vision
+7
Account Executive
Account Executive

Novara • Ontario (CA)

On-site
USD 86,000 - 107,000
Medical, Dental, Vision
401k with company match
No-cost mental health benefits
Head of Recruiting
Head of Recruiting

Novo • New York (NY), Northern (KY)

Hybrid
USD 180,000 - 260,000
Product & Privacy Counsel
Product & Privacy Counsel

Novacredit • New York (NY)

On-site
USD 223,300 - 281,400
Comprehensive medical, dental, and vision insurance
Company-sponsored 401k plan
Flexible PTO
Principal Manager, People Operations Business Partner
Principal Manager, People Operations Business Partner

Nava Public Benefit Corp. • Northern (KY)

Hybrid
USD 132,000 - 149,000
Health coverage
Insurance coverage
Time off
+15
Senior Software Engineer
Senior Software Engineer

Nova Credit Inc. • Tulsa (OK), Northern (KY)

Hybrid
USD 184,000 - 219,000
Principal Manager, People Operations Business PartnerRemote
Principal Manager, People Operations Business PartnerRemote

Nava Public Benefit Corp • United States

Remote
USD 132,000 - 149,000
Health coverage
Disability and life insurance
Generous paid time off & holidays
+2