AI-Native QA Engineer for LLM/Agent Quality

Newton Research

Massachusetts

On-site

USD 115,000 - 130,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Salary equity

Job summary

Newton Research is seeking a QA Engineer focused on AI-native quality in Boston/Needham, MA. You will own curated evals for agent behavior, validate prompts and skills, and ensure non-deterministic outputs are rigorously tested across the CI/CD pipeline.

You’ll collaborate with a release-readiness mindset, design guardrails, and drive regression testing for AI-driven features, with a strong emphasis on observability and reproducible results.

Qualifications

  • 4+ years in QA or test engineering on a complex web product, ideally B2B SaaS shipping frequently
  • Hands‑on experience building evals for LLM or agent products (datasets, rubrics, LLM‑as-judge, regression tracking)
  • Working knowledge of how agents work: prompting, skills, tool calling, context, RAG and orchestration
  • An automation-first instinct: you can show what you removed from a manual process
  • Python (or similar) for eval tooling and data checks; Playwright or similar for end‑to‑end tests
  • Statistical literacy: pass rates, variance and sample size when outputs are not deterministic
  • Daily user of AI coding and testing assistants, with the judgment to verify their output
  • Strong exploratory instincts and excellent written communication
  • Nice to have: adtech, martech or marketing analytics domain; SSO / SAML / OAuth flows; data‑connector testing; eval or observability tooling

Responsibilities

  • Own the eval suite for Newton's agents
  • Test skill and prompt changes before they ship
  • Cover the agent flows end to end
  • Handle non-determinism with rigor
  • Turn production and customer signal into evals
  • Probe AI-specific risk
  • Automate before you repeat
  • Design agentic QA workflows in CI
  • Keep human-judgment work sharp
  • Write bug reports that are machine- and human-readable

Skills

QA
Eval tooling
Automation
Python
Playwright
Statistics
Documentation
Written communication

Tools

Playwright
Python

Job description

Newton Research is seeking a QA Engineer focused on AI-native quality in Boston/Needham, MA. You will own curated evals for agent behavior, validate prompts and skills, and ensure non-deterministic outputs are rigorously tested across the CI/CD pipeline.

You’ll collaborate with a release-readiness mindset, design guardrails, and drive regression testing for AI-driven features, with a strong emphasis on observability and reproducible results.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

QA Engineer: AI-Native Quality, Equity & Impact
QA Engineer: AI-Native Quality, Equity & Impact

TestDevJobs • Boston (MA), Northern (KY)

Hybrid
USD 115,000 - 130,000
QA Engineer-AI Native Quality
QA Engineer-AI Native Quality

TestDevJobs • Boston (MA), Northern (KY)

Hybrid
USD 115,000 - 130,000
QA Engineer-AI Native Quality
QA Engineer-AI Native Quality

Newton Research • Massachusetts

On-site
USD 115,000 - 130,000
Salary equity
Lead AI QA Engineer for High-Stakes LLM Systems
Lead AI QA Engineer for High-Stakes LLM Systems

Novara • United States

On-site
USD 70,000 - 91,000
Medical
Dental
Vision
+6
QA AI Automation Engineer - LLMs & Multi-Agent Quality
QA AI Automation Engineer - LLMs & Multi-Agent Quality

Dynasty Financial Partners, LLC • Saint Petersburg (FL)

On-site
USD 120,000 - 150,000
AI-Native QA Engineer: Automated Testing for AI Systems
AI-Native QA Engineer: Automated Testing for AI Systems

IDR, Inc. • Rosemont (CO)

On-site
USD 120,000 - 160,000
Employee Stock Ownership Program
Full benefits
AI QA Automation Architect for LLM Systems
AI QA Automation Architect for LLM Systems

Dynasty Financial Partners • Town of Florida (NY), Northern (KY)

Hybrid
USD 120,000 - 150,000
QA AI Agent Developer — LLM-Driven Test Automation
QA AI Agent Developer — LLM-Driven Test Automation

Accord Technologies Inc • New Jersey

On-site
USD 90,000 - 120,000
Senior AI Quality Leader
Senior AI Quality Leader

LILT AI • City of Syracuse (NY)

Hybrid
USD 150,000 - 190,000
QA Engineer - Agentic Systems
QA Engineer - Agentic Systems

Meet Life Sciences • New York (NY)

On-site
USD 110,000 - 170,000