Lead AI QA Engineer for High-Stakes LLM Systems

Novara

United States

On-site

USD 70,000 - 91,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical
Dental
Vision
Flexible Spending Accounts
PTO and holidays
401k with company match
Life Insurance
Employee Assistance Programs
No-cost Mental Health Benefits

Job summary

Novara is seeking a hands-on QA/SDET specialist to own AI evaluation for non-deterministic systems. You will design frameworks, build test harnesses, and gate AI changes in CI/CD, partnering with product and engineering to raise quality for LLM-driven workflows.

You will mentor teams, contribute to API and end-to-end automation, and shape the QA playbook as agentic features spread across scrum teams, with strong focus on reliability, grounding, and latency.

Qualifications

  • 6+ years in QA, SDET, or test automation, with real production automation shipping
  • Hands-on experience testing LLM-based or agentic systems: building evals, working with LLM-as-judge patterns, prompt regression testing, or agent trajectory analysis
  • Prior experience in a shift-left, embedded QA model
  • Comfort with at least one modern automation stack (Playwright, Cypress, or similar) and a typed language, TypeScript preferred, Python fine
  • Deep experience with test frameworks such as vitest, jest, or pytest, and comfort building custom test harnesses rather than only running off-the-shelf suites
  • API-first testing mindset, including REST and Postman or equivalent
  • Fluency in HTTP-level API testing, including recording proxies and observing service-to-service traffic
  • Working knowledge of CI/CD pipelines, GitHub Actions a plus, and how to plug evals into them
  • Familiarity with cloud secret managers and disciplined handling of sensitive test data in restore-from-prod environments
  • Ability to reason clearly about probabilistic systems: variance, sample sizes, confidence, and when a flaky result is signal rather than noise

Responsibilities

  • Design and maintain evaluation frameworks for LLM outputs and agentic workflows, including regression suites, golden datasets, and scoring rubrics
  • Build test harnesses that catch hallucinations, tool-calling failures, prompt regressions, and unsafe or off-policy behavior before they reach production
  • Define measurable quality criteria for agent reliability: task completion, factual grounding, latency, cost, and reasoning quality
  • Integrate evaluation runs into CI/CD so model, prompt, and agent changes are gated the same way code changes are
  • Partner with engineers on observability and tracing for agent runs, so failures are diagnosable rather than mysterious
  • Contribute to conventional API and end-to-end automation where AI features sit inside larger product flows
  • Help shape the team's shared playbook for testing AI features, and coach other QA engineers as agentic work spreads across scrum teams

Skills

QA automation
LLM testing
Agentic systems
Playwright
Cypress
Vitest/Jest/Pytest
TypeScript
REST API testing
CI/CD
GitHub Actions
Observability/tracing
Data handling in prod

Tools

Playwright
Cypress
Postman
Datadog
Jira/Xray
AWS

Job description

Novara is seeking a hands-on QA/SDET specialist to own AI evaluation for non-deterministic systems. You will design frameworks, build test harnesses, and gate AI changes in CI/CD, partnering with product and engineering to raise quality for LLM-driven workflows.

You will mentor teams, contribute to API and end-to-end automation, and shape the QA playbook as agentic features spread across scrum teams, with strong focus on reliability, grounding, and latency.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM QA Engineer: AI Testing & Evaluation
LLM QA Engineer: AI Testing & Evaluation

Codefeast • United States

On-site
USD 90,000 - 140,000
AI QA Automation Architect for LLM Systems
AI QA Automation Architect for LLM Systems

Dynasty Financial Partners • Town of Florida (NY), Northern (KY)

Hybrid
USD 120,000 - 150,000
AI-Native QA Engineer for LLM/Agent Quality
AI-Native QA Engineer for LLM/Agent Quality

Newton Research • Massachusetts

On-site
USD 115,000 - 130,000
Salary equity
Lead QA Engineer for AI & LLM Workflows (Remote)
Lead QA Engineer for AI & LLM Workflows (Remote)

Vibehackers • Northern (KY)

Hybrid
USD 150,000 - 190,000
Generous annual bonus opportunity
401(k) with employer match
Medical Insurance
+2
QA AI Automation Engineer - LLMs & Multi-Agent Quality
QA AI Automation Engineer - LLMs & Multi-Agent Quality

Dynasty Financial Partners, LLC • Saint Petersburg (FL)

On-site
USD 120,000 - 150,000
LLM QA Engineer - Automated Testing & Prompting
LLM QA Engineer - Automated Testing & Prompting

ICE • Atlanta (GA)

On-site
USD 80,000 - 130,000
QA Engineer - Agentic Systems
QA Engineer - Agentic Systems

Meet Life Sciences • New York (NY)

On-site
USD 110,000 - 170,000
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
Senior AI Quality Leader
Senior AI Quality Leader

LILT AI • City of Syracuse (NY)

Hybrid
USD 150,000 - 190,000
Senior AI Engineer - Production LLM EvalOps
Senior AI Engineer - Production LLM EvalOps

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000